Site Reliability Engineer Jobs

Engineers who ensure cloud services are reliable, scalable, and efficient. A critical role in maintaining uptime and performance in cloud environments.

Open roles
16
Salary range
£46k – £100k
Hiring companies
10

Site Reliability Engineers (SREs) are the backbone of cloud operations, ensuring that applications and services run smoothly and efficiently. They work closely with development teams to build and maintain robust, scalable systems that can handle high traffic and complex workloads. SREs are in high demand across a range of industries, from tech startups to large enterprises, and their role is crucial in maintaining the reliability and performance of cloud-based applications.

What the role does

Inside the role of a Site Reliability Engineer

A typical week for an SRE is a mix of proactive system maintenance, incident response, and collaboration with development teams.

  1. 01
    Monitor system performance and health metrics.
  2. 02
    Respond to and resolve incidents and outages.
  3. 03
    Implement and optimise automation scripts and tools.
  4. 04
    Collaborate with developers on system design and improvements.
  5. 05
    Conduct post-incident reviews and document findings.
  6. 06
    Participate in on-call rotations and provide 24/7 support when needed.
Salary on the board

£46k – £100k

Based on advertised midpoints across the 88 UK listings priced in pounds in the last 12 months. Base salary only.

Salary visibility
47% of listings advertise a salary — up from 0% the year before.
By seniority
£k base
Mid
44
92
19 jobs
Senior
60
95
23 jobs
Lead
64
100
21 jobs
Skills & tools

What hiring managers ask for

% of 57 listings posted in the last 12 months that mention each skill, extracted from job descriptions.

Terraform
68%
CI/CD
63%
Python
54%
AWS
51%
Linux
44%
Observability
42%
Azure
40%
Kubernetes
39%
Grafana
37%
Prometheus
37%
Automation
33%
Monitoring
32%
Career ladder

From Junior to Principal

A typical UK progression for site reliability engineers. Years are guidance — strong people move faster, and many senior folks sidestep into research, product or management.

  1. Level 1

    Junior Site Reliability Engineer

    0–2 yrs

    Assist in monitoring and maintaining cloud infrastructure, with a focus on learning and supporting more experienced team members.

  2. Level 2

    Site Reliability Engineer

    2–5 yrs

    Own the reliability and performance of specific systems, implementing automation and optimisation strategies.

  3. Level 3

    Senior Site Reliability Engineer

    5–8 yrs

    Lead the design and implementation of complex cloud architectures, mentor junior engineers, and drive reliability initiatives.

  4. Level 4

    Principal Site Reliability Engineer

    8+ yrs

    Strategise and oversee the reliability and scalability of the entire cloud infrastructure, influencing company-wide practices and policies.

Pathway

How to become a Site Reliability Engineer

There's no single route, but most people follow some version of these steps.

  1. 1

    Learn the Basics

    Start with foundational knowledge in cloud platforms, scripting, and system administration.

  2. 2

    Gain Practical Experience

    Work on real-world projects, contributing to monitoring, automation, and incident response.

  3. 3

    Specialise in SRE Practices

    Deepen your expertise in reliability engineering, including capacity planning and disaster recovery.

  4. 4

    Lead Projects and Teams

    Take on leadership roles, managing teams and driving large-scale reliability initiatives.

  5. 5

    Influence Company Strategy

    Shape the company's cloud strategy and best practices, contributing to long-term reliability and efficiency.

Live jobs

16 live roles

See all 16 roles →

Site Reliability Engineer

This role involves ensuring the reliability, scalability, and performance of large-scale, business-critical platforms in a fully remote environment. The Site Reliability Engineer will diagnose and resolve complex production issues, implement automation to reduce manual toil, and use Infrastructure as Code to manage environments. Close collaboration with development and DevOps teams is required to balance new feature delivery with system stability and security.

Bristow Holland Ltd Manchester, United Kingdom £55,000 – £60,000 pa

Senior Site Reliability Engineer — Token Factory (Inference Platform)

Design and maintain telemetry pipelines for metrics, logs, and traces at scale, while optimizing Kubernetes and Terraform configurations for GPU-intensive inference workloads. Focus on building self-healing, observable systems that ensure high reliability and performance across a distributed AI infrastructure. Collaborate with engineering teams to harden request routing, autoscaling, and incident response mechanisms for large-scale model deployment.

Nebius United Kingdom
Hiring locations

Where this role is hiring

The locations with the most live listings for this role today.

FAQs

Common questions

  • Essential skills include strong knowledge of cloud platforms, scripting, system administration, and automation tools. Familiarity with monitoring and incident response is also crucial.

  • SREs collaborate closely with developers to ensure that applications are designed for reliability and scalability. They provide feedback on system design and help implement automation and monitoring solutions.

  • SREs often work in fast-paced, collaborative environments. They may be part of on-call rotations and need to be available to respond to incidents at any time.

  • Advancement involves gaining experience, specialising in advanced SRE practices, and taking on leadership roles. Continuous learning and staying updated with the latest cloud technologies are also important.

  • Salary ranges can vary widely based on experience, location, and company size. For more detailed information, please refer to the salary section on this page.

Hiring site reliability engineers?

Post your role in 90 seconds and reach the specialist audience that already reads this page.