Latest Cloud Infrastructure Jobs

Data Center IT Support Manager

This role involves managing IT support operations for a high-performance data center focused on AI infrastructure, leading teams of engineers, overseeing GPU cluster maintenance, incident resolution, and service-level performance. The position requires hands-on technical leadership, process improvement, and coordination with infrastructure and vendor teams. It plays a critical part in ensuring reliable, scalable operations within a fast-moving AI cloud environment.

Nebius United Kingdom

Product Manager - Security

This role involves defining and executing the product strategy for Nebius' security offerings, ensuring compliance and certifications are effectively communicated to customers and stakeholders. The Product Manager will lead external security communications, drive certification roadmaps, and collaborate with engineering and security teams to position security as a competitive advantage. The role requires translating technical security posture into customer-facing value while working across compliance, product, and go-to-market functions.

Nebius United Kingdom
Remote Permanent
CrowdStrike logo

Engineering Manager - Core Infrastructure

This role involves leading a team of engineers to build and maintain core infrastructure platforms at scale, focusing on infrastructure-as-code, cloud cost optimization, and high-availability systems. The Engineering Manager will drive technical strategy, enforce deployment standards across multi-cloud environments, and integrate AI-driven workflows to enhance platform efficiency. Collaboration through design reviews, roadmap planning, and cross-team technical leadership is central to the position.

CrowdStrike United States £140,000 – £215,000 pa
Remote
CrowdStrike logo

Engineering Manager - Core Infrastructure

This role involves leading a team of engineers to build and maintain core infrastructure platforms at a large-scale cybersecurity company. The focus is on self-service infrastructure-as-code, cloud cost optimization, and enforcing deployment standards across a multi-cloud environment. The manager will also drive technical strategy, mentor senior engineers, and integrate AI-driven improvements into workflows.

CrowdStrike US$140,000 – US$215,000 pa
Remote

Systems Integration Engineer

Design and implement integrations between core business systems including ERP, HRIS, CRM, and identity providers, ensuring reliable data flow and compliance. Build automated joiner/mover/leaver workflows and support audit requirements. Work with Palantir Foundry and API-based platforms to establish scalable integration patterns across a growing enterprise environment.

NScale London, United Kingdom
On-site Permanent

Staff HPC Systems Software Engineer

This role involves defining the technical direction for a core HPC systems domain within a cloud-native AI infrastructure platform. You'll architect and evolve Slurm-based HPC services, drive cross-team integration with Kubernetes-adjacent systems and infrastructure APIs, and establish reusable patterns for automation, reliability, and operational supportability. The position requires deep systems expertise in HPC, GPU scheduling, and production-scale software engineering to ensure robust, maintainable platform capabilities.

NScale United Kingdom

Senior AWS Site Reliability Engineer

This role involves ensuring high availability and performance of large-scale distributed systems by building automated solutions and driving reliability across cloud and on-prem environments. You'll collaborate with development teams to improve deployment practices, lead incident response, and optimise infrastructure through observability and performance tuning. The focus is on enhancing system resilience, scalability, and delivery speed using modern DevOps practices and cloud-native technologies.

Spectrum IT Recruitment London, United Kingdom £60,000 – £70,000 pa
Hybrid Permanent
NVIDIA logo

Senior HPC AI Cluster Engineer

Design and maintain large-scale HPC/AI clusters with a focus on system automation, performance tuning, and infrastructure scalability. Develop CI/CD pipelines and monitoring solutions while collaborating with researchers and engineers to optimize workflows for accelerated computing. Support R&D through proof-of-concepts and contribute to standard methodologies for deploying GPU-accelerated systems.

NVIDIA Germany
Remote Permanent
NVIDIA logo

Senior HPC AI Cluster Engineer

Design, deploy, and maintain large-scale HPC and AI clusters with a focus on accelerated computing and GPU workloads. Develop automation tooling for infrastructure deployment, monitoring, and self-service resource consumption. Troubleshoot across bare metal, OS, software stack, and application layers while supporting R&D through proof-of-concept initiatives and performance optimization.

NVIDIA
Remote Permanent

VDI Engineer

This role involves supporting and optimising a large-scale enterprise VDI and End-User Computing (EUC) environment using VMware Horizon, App Volumes, DEM, and Instant Clones. The engineer will manage core Windows infrastructure including Active Directory, DNS, DHCP, and Group Policy, while automating tasks with PowerShell and performing root cause analysis on complex issues. Collaboration with security, cloud, and project teams is key to delivering platform improvements and lifecycle management.

Synapri Glasgow, City Of Glasgow, G2 1AL, United Kingdom £600 – £650 pd

Staff Backend Engineer - Alerting | UK |

This role involves designing, building, and maintaining highly scalable distributed systems for Grafana's alerting infrastructure, with a focus on performance, reliability, and cross-platform consistency. You'll work deeply in backend systems handling Prometheus and Grafana-managed alerts, contribute to open-source projects, and operate services at global scale. The position emphasizes technical ownership, mentorship, and collaboration in a fully remote, innovation-driven environment.

Grafana Labs United Kingdom £103,958 – £124,750 pa
Remote Permanent

Staff Software Engineer - Databases SRE | UK |

This role involves ensuring the reliability and scalability of Grafana Cloud's core database products (Mimir, Loki, Tempo, Pyroscope) for high-SLA customers. You'll work embedded within engineering squads to design SLOs, automate reliability practices, lead incident response, and improve observability in multi-tenant cloud environments. The position emphasizes technical leadership, production excellence, and influencing product design for operability across AWS, GCP, and Azure.

Grafana Labs United Kingdom £103,958 – £124,750 pa
Remote Permanent

DevOps Platform Engineer

This role involves hands-on management and scaling of Kubernetes clusters, database administration, and automation of infrastructure. The engineer will ensure high availability, security, and performance of critical systems while supporting global engineering teams.

Context Recruitment London, United Kingdom £90,000 – £100,000 pa

MLOps Engineer

This role involves owning and developing the MLOps infrastructure for a growing AI platform, building and maintaining orchestration pipelines, deploying ML workloads on Kubernetes, and implementing CI/CD and monitoring systems. You will work closely with Data Scientists and ML Engineers to scale deployments and support real-world impact.

Harnham - Data and Analytics Recruitment London, United Kingdom £75,000 – £85,000 pa

IT Delivery Manager

The IT Delivery Manager will lead multiple IT projects, including infrastructure, cloud, applications, and security. Responsibilities include project planning, documentation, resource management, and stakeholder communication. The role requires strong technical awareness, leadership, and ITIL-aligned change management experience.

Metaskil Limited Hatfield, United Kingdom £50,000 pa