(Senior) Infrastructure Engineer

NScale
London, W1U 7ED, United Kingdom
2 months ago
Applications closed

Related Jobs

View all jobs

Data Centre Technician , Data Center Operations

Amazon Didcot, United Kingdom
On-site

Solutions Architect, Financial Services, Insurance

Amazon London, United Kingdom
On-site

Senior Software Engineer II, Kora Compute

Confluent United Kingdom
On-site

GenAI/ML Specialist Solutions Architect, AWS Global Sales (AGS)

Amazon London, United Kingdom
On-site

Senior Technical Infrastructure Program Manager, Networking

Amazon London, United Kingdom
On-site Clearance Required

Senior Customer Success Specialist, Connect Specialty Sales

Amazon London, United Kingdom
On-site
Job Type
Permanent
Work Pattern
Full-time
Work Location
On-site
Seniority
Senior
Education
Degree
Posted
28 May 2026 (2 months ago)

Benefits

On-call rotations Incident response activities

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.

About the Role

We’re hiring anInfrastructure Engineer to design, implement, operate, and continuously improve the infrastructure platforms that support both internal and customer-facing services at Nscale.

This role sits within theOperational Engineering team inEngineering, where you’ll work across the infrastructure stack below the hypervisor with a focus onLinux,OpenStack, storage systems, Proxmox, DNS, DHCP, and infrastructure automation. You’ll collaborate closely with internal teams to ensure infrastructure meets performance, availability, and security requirements, while also serving as a3rd/4th line escalation point for complex issues.

Your work will directly support the reliability, scalability, automation, and security of the platforms that power Nscale’s GPU cloud. This is a high-impact role for someone who wants to shape core infrastructure, improve operational excellence, and bring deep technical expertise to both delivery and ongoing evolution of critical systems.

What you'll be doing

Infrastructure Design & Operations

  • Implement infrastructure components that underpin internal and customer-facing services.
  • Operate critical infrastructure layers below the hypervisor with a focus on stability and performance.
  • Maintain essential services such asDNS, DHCP, and configuration management tooling.
  • Design scalable and resilient infrastructure platforms acrossOpenStack, Proxmox, Ceph, and core supporting services.

Automation & Continuous Improvement

  • Improve automation for provisioning, monitoring, patching, and recovery.
  • Use infrastructure-as-code and configuration management tools to standardise operations.
  • Drive continuous improvement across infrastructure reliability, scalability, and operational efficiency.
  • Support repeatable and maintainable platform operations through automation-first approaches.

Incident Management & Escalation

  • Act as a3rd/4th line escalation point for complex infrastructure issues.
  • Partner with support teams to resolve incidents and restore services effectively.
  • Investigate root causes of infrastructure problems and contribute to long-term fixes.
  • Participate inon-call rotations and incident response activities for critical infrastructure.

Cross-Functional Collaboration & Technical Guidance

  • Collaborate with internal teams to ensure solutions meetperformance, availability, and security requirements.
  • Contribute to infrastructure roadmap planning, includingcapacity management andperformance tuning.
  • Introduce new technologies that strengthen the infrastructure stack over time.
  • Provide technical expertise topre-sales and other groups on infrastructure capabilities and best practices.

Standards, Security & Compliance

  • Ensure infrastructure platforms adhere to compliance, security, and operational standards.
  • Apply best practices to the operation and evolution of infrastructure services.
  • Support secure and well-governed platform delivery across the environments you own.

KPIs

  • Infrastructure availability and resilience
  • Automation coverage for provisioning, patching, monitoring, and recovery
  • Complex incident resolution and root cause remediation
  • Capacity management and performance tuning effectiveness

About You

  • Strong Python and Bash skills
  • Strong troubleshooting experience withLinux and services running on Linux
  • Experience working withCeph and core infrastructure services
  • Nice to have experience deploying, managing, upgrading, and operating largeOpenStack clusters
  • Experience deploying, managing, and automatingProxmox
  • Knowledge ofDNS, DHCP, and configuration management in production environments
  • Ability to operate and improve infrastructure with a focus onavailability, scalability, automation, and security
  • Experience handling complex infrastructure issues in an escalation capacity
  • Ability to work effectively with internal teams and provide technical input across the organisation
  • Nice to have knowledge of Ironic andNeutron/OVN/OVS is a plus

What we can offer you

At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.

Highly competitive compensation package (base + bonus + equity), with performance reviews every 12 months. 🚀

Join one of the fastest-growing AI infrastructure companies — your chance to directly shape how global AI capacity is planned and deployed. ✨

Expect a dynamic progression plan tailored to your ambitions. Grow by leading critical cross-functional initiatives and shaping capital strategy — always with our full support.

Human-First Flexibility: We treat you as humans first. 🫶🏽 Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments.

Equal Opportunities Statement

We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.

If there’s anything we can do to accommodate your specific situation, please let us know.

The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Industry Insights

Discover insightful articles, industry insights, expert tips, and curated resources.