Latest Cloud Infrastructure Jobs

Senior Site Reliability Engineer (SRE)

Ensure fault-tolerance, scalability, and continuous operation of cloud infrastructure by solving complex engineering challenges across compute, storage, and networking. Work with cutting-edge technologies to optimize large-scale GPU orchestration and AI/ML deployment. Collaborate in a global, fast-paced environment focused on building a full-stack AI cloud platform.

Nebius United Kingdom

Senior Software Engineer (Managed PostgreSQL)

This role involves developing and optimizing a managed PostgreSQL service tailored for AI workloads, with a focus on vector search, migration tooling, and operational reliability. The engineer will work across the full stack—from Postgres internals and Go-based control plane systems to customer-facing performance escalations. Key responsibilities include enhancing database performance, enabling seamless migrations from other cloud providers, and integrating with the broader Nebius AI Cloud platform.

Nebius United Kingdom

Deployment Engineering Director, Systems Engineering

Lead and scale a high-performing engineering team responsible for core platform systems, driving multi-quarter initiatives to improve delivery velocity, system reliability, and product quality through systems validation. Partner with product and engineering leaders to translate ambiguous, high-impact challenges into clear execution plans, ensuring operational excellence across complex, interdependent infrastructure and developer tooling.

NScale United Kingdom

Reliability Engineer SME (M&E)

This role involves providing expert mechanical and electrical engineering support across Nscale's EMEA data centre estate, with a focus on reliability, resilience, and performance for high-density AI compute infrastructure. The engineer will lead technical escalations, design reviews, change approvals, incident resolution, and RCA, while ensuring compliance with engineering standards across both owned and third-party sites. Key responsibilities include supporting power and cooling systems, conducting audits, and improving energy efficiency in mission-critical environments.

NScale London, United Kingdom

Senior Cyber Security Analyst

This role involves leading the detection, investigation, and response to cyber security incidents across the university’s systems and cloud environments. The analyst will strengthen security controls, manage vulnerabilities, and embed security into digital projects and operations. Key responsibilities include improving threat detection, supporting compliance, and leading phishing awareness initiatives within a collaborative security team.

De Montfort University Leicester Leicester, LE1 5YA, United Kingdom £44,746 – £56,535 pa

Senior DevOps Systems Administrator

This role involves designing and deploying hybrid cloud infrastructure, with a focus on AWS and private cloud solutions. The candidate will lead automation initiatives using Terraform and Ansible, manage Kubernetes (EKS) and virtualisation platforms, and build CI/CD pipelines using GitHub Actions. They will also mentor team members, ensure system security and reliability, and contribute to incident response through on-call duties.

Rise Technical Recruitment Guildford, United Kingdom £60,000 – £65,000 pa

Platform Engineer | Kubernetes | AWS/Azure | HealthTech

This role involves managing and evolving a self-managed Kubernetes platform across AWS, Azure, and on-prem environments to support critical health services. The engineer will focus on building secure, resilient, and high-performance infrastructure at scale. The position offers technical ownership and impact in a healthcare technology setting.

Opus Recruitment Solutions Binfield, Berkshire, United Kingdom £65,000 – £75,000 pa

Technology Director

The Technology Director will lead the technology strategy, focusing on software development, infrastructure, data, and AI. Key responsibilities include identifying commercial impacts, improving systems, and leading the development function. The role involves shaping the technology roadmap and ensuring the business scales successfully.

RaptorTech Recruitment Nr11Bj, United Kingdom £70,000 – £80,000 pa

Staff Software Engineer - Databases, Tempo | United Kingdom |

Design and evolve core components of Tempo, Grafana's open-source distributed tracing backend, focusing on scalability, performance, and platform extensibility. Lead architectural initiatives, improve query and ingestion systems, and enable AI-driven observability workflows. Work in a fully remote, asynchronous environment with deep operational ownership and a strong open-source ethos.

Grafana Labs United Kingdom £103,958 – £124,750 pa
Remote Permanent

Senior Site Reliability Engineer

This role involves ensuring the reliability, performance, and scalability of large-scale distributed systems in a cloud and on-premises environment. You'll develop automation tools, manage Kubernetes infrastructure, and work closely with development teams to improve service quality and deployment practices. The focus is on observability, incident management, and optimising platform health using modern DevOps tooling and cloud-native technologies.

Spectrum IT Recruitment Southampton, Hampshire, United Kingdom £60,000 – £70,000 pa

Senior Backend Engineer - Databases Pyroscope | UK |

Design, build, and operate core components of Pyroscope, a distributed profiling database, focusing on scalability, cost efficiency, and deep integration with the broader Grafana stack. Lead end-to-end projects such as adaptive profiling, trace-to-profile correlation, and automation for operational excellence. Engage directly with customers and the open-source community to shape product direction in a PM-light engineering culture.

Grafana Labs United Kingdom £91,755 – £110,106 pa
Remote Permanent

Senior Data Engineer

This role involves architecting and evolving scalable data infrastructure to support decision-making in a fast-paced, analytics-focused environment at the intersection of sport and technology. You will lead technical design, mentor engineers, and work closely with cross-functional teams to deliver advanced data solutions and maintain data quality and security.

Tenth Revolution Group London, United Kingdom £65,000 – £75,000 pa
Hybrid Permanent
NVIDIA logo

Senior MLOps Engineer - DSX Enablement

Develop and optimize full-stack MLOps pipelines for large-scale AI and ML workloads, with a focus on distributed training, inference performance, and integration with NVIDIA's hardware and cloud platforms. Collaborate with internal teams and external customers to solve complex system-level challenges across hardware, networking, and software stacks. Contribute to open-source tools and reference architectures to advance scalable AI infrastructure.

NVIDIA Germany PLN 292,500 – PLN 650,000 pa
NVIDIA logo

Senior MLOps Engineer - DSX Enablement

Develop and optimize full-stack AI/ML systems on NVIDIA's cloud platforms, focusing on MLOps pipelines, distributed training, and inference performance. Serve as a technical advisor for internal and external customers, solving complex production issues and building open-source tools to scale AI workloads. Collaborate with infrastructure teams to enhance frameworks and support new hardware integration.

NVIDIA PLN 292,500 – PLN 650,000 pa
Databricks logo

Staff Software Engineer - Backend

Design and build scalable backend systems for a cloud-based data and AI platform, focusing on distributed systems, reliability, and performance at scale. Work across the full development lifecycle to improve data storage, access, and platform stability. Contribute to foundational components of the Lakehouse architecture while collaborating with cross-functional teams.

Databricks London, United Kingdom
On-site Permanent