Cloud Operations - Service Reliability Engineer

A&O Shearman
Bt132Bj, United Kingdom
Today
Posted
3 Aug 2026 (Today)
What you will do

The Service Reliability Engineer is accountable for improving the reliability, observability and operational resilience of cloud-hosted services. The role focuses on monitoring, early issue identification, cloud engineering and automation, using tools and practices such as Bicep, Azure DevOps, GitHub and Ansible to support consistent, repeatable and well-governed service operation.
  • Monitoring, observability and alerting across cloud infrastructure, platform services and supported application environments;
  • Cloud engineering and automation, including Infrastructure as Code, deployment pipelines, configuration management and standards-led delivery; and
  • Promoting service resiliency through proactive issue identification, operational insight, automation and continuous improvement.
The role involves:
  • Supporting service reliability, observability and cloud engineering across the following areas:
    • Monitoring, observability and alerting for cloud-hosted services, including infrastructure health, service availability, performance signals and operational events - Essential;
    • Azure public cloud engineering, including IaaS, PaaS, networking, identity, RBAC and platform diagnostics - Essential;
    • Infrastructure as Code and automation using Bicep, Azure DevOps pipelines and GitHub-based source control and collaboration - Essential;
    • Configuration management and standards automation using Ansible or equivalent tooling - Preferred;
    • Experience of using or implementing monitoring solutions using Elastic - Preferred;
    • Operational reporting, issue trend analysis and the development of actionable dashboards to support service improvement - Preferred;
    • Experience of working across a broad range of systems, technologies and internal support teams - Preferred.
  • Ensuring that monitoring and operational insight are effectively designed, implemented and understood so that services can be supported, improved and made more resilient.
  • Providing subject matter expertise in cloud operations, observability, automation and reliability engineering practices.
  • Working globally across cloud-hosted services and platform capabilities, independent of location.
  • Support the firm's environmental goals and initiatives.

Monitoring, Reliability and Cloud Engineering

  • Works with internal technology teams to improve end-to-end observability for supported services, including:
    • Monitoring coverage for infrastructure, platform services and application components;
    • Actionable alerting that supports early identification of degradation, failure or operational risk;
    • Dashboards and reporting that help teams understand service health, trends and recurring issues;
    • Cloud engineering practices that use Bicep, Azure DevOps, GitHub and Ansible to deliver consistent and repeatable change; and
    • Operational standards that improve service resilience and reduce manual support effort.
  • Maintains appropriate documentation, including monitoring standards, known issues, operational patterns, troubleshooting guidance and support handbooks.

Service Delivery

  • Identify, diagnose and support resolution of incidents and problems by interpreting monitoring signals, operational telemetry and service behaviour.
  • Work with service-owning teams to improve the quality, relevance and routing of alerts so that operational issues can be detected and acted upon quickly.
  • Contribute to root cause analysis, problem management and continuous improvement activity by identifying recurring patterns, gaps in observability and opportunities for automation.

Build and Implementation

  • Provide specialist guidance to teams adopting cloud engineering patterns, Infrastructure as Code, deployment pipelines and automated configuration management.
  • Support implementation of monitoring and automation standards across new and existing services.
  • Ensure that operational documentation, handover materials and support guidance are created and are suitable for BAU operation.

Risk Management

  • Identify operational, reliability and supportability risks arising from gaps in monitoring, alerting, automation or cloud platform standards.
  • Refer to domain experts for guidance on specialised areas such as architecture, security, networking, database platforms and application design.
  • Participate in recovery, resilience and operational readiness activities to help prove that services can be supported effectively.

Quality, Methods & Tools

  • Strive for improvements to processes by promoting standardised patterns, automated controls, repeatable engineering practices and effective use of industry best practice.
  • Advocate for the use of source control, pipeline-based delivery, Infrastructure as Code and configuration management to improve quality, auditability and operational reliability.

What you will have

Business Competencies

  • Strong analytical and problem-solving skills, with a logical approach to issue identification, diagnosis and service improvement.
  • Technically curious, with an enthusiasm for understanding a broad set of systems, technologies and operational domains.
  • Ability to interpret monitoring data, identify patterns and translate operational insight into meaningful improvement activity.
  • Ability to make sound decisions under pressure and support effective incident response.
  • Strong commitment to service reliability, operational resilience and excellent customer service.
  • Commercial acumen, including an understanding of IT service costs, cloud consumption and how technology adds value to the business.
  • Ability to promote technical standards, automation and reliability practices using clear, business-friendly language.
  • Personal credibility; highly self-motivated self-starter who will undertake all activities to the highest professional standards.
  • Excellent communication skills, both orally and written.
  • Ability to operate within a wider team where there may be ambiguity and conflicting priorities.
  • Ability to build effective working relationships across a diverse set of internal teams and influence the adoption of monitoring, automation and cloud engineering standards.
  • Experience of working in a global environment across international locations with an appreciation of multiple cultures.

Knowledge

Practical knowledge of SRE principles, observability, incident response, problem management and operational resi

Related Jobs

View all jobs

Technical Lead AWS Operations Engineer

Spectrum IT Recruitment Southampton, Hampshire, SO19 8NJ, United Kingdom

Google Cloud Application Support Engineer

Summer Browning Associates London, United Kingdom
£1 pd

Cloud Engineer

Yunex Limited Broadstone, Dorset, DT11 0QJ, United Kingdom
Hybrid

Network Engineer- Cloud , London)

CrowdStrike London, United Kingdom
Hybrid

AWS Infrastructure Engineer (Night Shift)

Spectrum IT Recruitment United Kingdom
£45,000 – £60,000 pa

Technical Account Manager, Enterprise Support - TMEGS

Amazon London, United Kingdom
On-site

Industry Insights

Discover insightful articles, industry insights, expert tips, and curated resources.