Site Reliability Engineer
- Remote job
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer based in United Kingdom.
This is a senior reliability engineering role focused on defining and advancing the infrastructure standards behind a globally scaled, AI-native platform.
You will take ownership of reliability strategy across production infrastructure, with a particular focus on AWS, Kubernetes, event-driven systems, and AI agent workloads.
The role combines deep hands-on engineering with architectural leadership, incident management, observability, and technical mentorship.
You will design systems that remain resilient under increasing transaction volumes while establishing measurable standards for reliability across engineering teams.
A key part of the role will be evolving synchronous architectures toward durable asynchronous communication and strengthening the platform through resilience testing and chaos engineering.
You will also help shape how AI-assisted tools are used for automation, incident analysis, runbooks, and root-cause investigations.
Success means becoming the trusted technical authority for complex reliability decisions while creating practices that make reliability scalable across the organization.
- Define and implement the reliability strategy across the platform, including SLOs, SLIs, error budgets, incident practices, and reliability standards adopted by engineering teams.
- Drive major architectural decisions as infrastructure evolves, evaluating technologies and designing systems that remain scalable, resilient, observable, and maintainable.
- Design and own event-driven communication and messaging infrastructure, including the transition from synchronous patterns to durable asynchronous architectures.
- Manage and evolve cloud infrastructure on AWS, using Infrastructure as Code to automate provisioning, configuration, deployment, and operational processes.
- Ensure Kubernetes and containerized workloads scale reliably as transaction volumes and AI workloads increase.
- Build and maintain comprehensive observability through monitoring, dashboards, alerting, application performance monitoring, and distributed tracing.
- Serve as the senior escalation point for complex production incidents, leading incident response, root-cause investigations, and blameless postmortems.
- Turn incident findings into permanent improvements through architectural changes, automation, operational controls, and resilience patterns.
- Establish a continuous chaos engineering and resilience testing practice through fault injection, game days, and controlled failure experiments.
- Mentor senior and mid-level engineers while raising the technical bar for reliability engineering and influencing engineering practices across teams.
- Use AI-assisted tooling for automation, runbooks, incident analysis, and root-cause investigations, while helping establish effective AI-enabled engineering practices.
- Within the first 6–12 months, establish the platform reliability strategy, lead at least one major architectural evolution, and drive adoption of the SLO and error-budget framework across engineering teams.
- Extensive experience in Site Reliability Engineering, Platform Engineering, DevOps, or a closely related discipline, with demonstrated ownership of production-scale systems.
- Deep expertise in event-driven architecture and messaging systems such as Kafka, NATS, or RabbitMQ, including at-least-once delivery, consumer groups, dead-letter queues, backpressure, and migrations from synchronous to asynchronous architectures.
- Strong AWS expertise across services such as EC2, VPC, IAM, S3, and RDS, combined with solid networking fundamentals.
- Hands-on Infrastructure as Code experience using Terraform, Pulumi, or similar tools, with infrastructure managed through version-controlled workflows and code reviews.
- Strong production experience with Kubernetes and Docker, including container lifecycle management, resource limits, health checks, and orchestration at scale.
- Proven observability expertise using Datadog or equivalent platforms, including dashboards, monitoring, APM, distributed tracing, and alerting.
- Demonstrated experience defining and operating SLOs, SLIs, and error budgets across multiple services.
- Hands-on experience with chaos engineering, fault injection, game days, or resilience experiments using tools such as Gremlin, Chaos Mesh, AWS FIS, or similar technologies.
- Strong distributed systems debugging skills, with experience diagnosing asynchronous workflows, cascading failures, and complex production incidents.
- Ability to code for automation and engineering tooling using Go, Python, or a similar programming language.
- Solid database knowledge across SQL and NoSQL technologies, particularly PostgreSQL, MongoDB, and Redis, including indexing, replication, and performance optimization.
- Proven technical leadership experience, including setting reliability standards, influencing architecture across teams, and mentoring engineers.
- Advanced written and spoken English communication skills.
- Experience with AI or MLOps infrastructure, including model serving, LLM inference, GPU/resource management, or AI agent observability, is highly advantageous.
- Familiarity with multi-tenant container platforms and customer workload infrastructure is a plus.
- Experience with data pipelines and orchestration tools such as Airflow or Prefect, and data platforms such as Databricks, Snowflake, or BigQuery, is beneficial.
- Familiarity with incident management platforms such as PagerDuty, Opsgenie, or incident.io is an advantage.
- Experience in the payments industry is preferred.
- Additional experience with ECS, s6-overlay, AI agent frameworks, or Spanish proficiency is a plus.
- Competitive compensation.
- Fully remote working environment with the flexibility to work from different locations.
- One-time home office allowance to help create an effective workspace.
- Company-provided work equipment.
- Stock options.
- Health plan available wherever you are.
- Flexible days off.
- Access to language, professional, and personal development courses.
- Opportunity to work on globally scaled infrastructure supporting complex payment and AI workloads.
- Significant technical ownership and influence over reliability strategy, architecture, and engineering standards.
- Collaborative international environment with opportunities to mentor engineers and shape organization-wide engineering practices.
Requirements
Benefits
£50k - £65k per annum
...Calling all SREs: Help solve complex monitoring and automation challenges! Are you a Site Reliability Engineer or a Software Developer looking to join a team and help solve complex monitoring, service management, automation, and capacity management problems.? As part of...SuggestedPermanent£50k - £70k per annum
...A Software Engineer possesses a unique skill set that synergises well with Site Reliability Engineering. With a strong foundation in software development, valuable expertise are bought to the table, enabling contribution to innovative solutions for complex monitoring, automation...SuggestedPermanent£45k - £65k per annum
...'t the right answer). That was until today, when I met a company who actually white boarded their vision for a brand new Site Reliability Engineering function they're building from scratch. They talked about production infrastructure, optimisation, automation and focusing...SuggestedPermanent£60k - £70k per annum
...Are you a seasoned Site reliability Engineer looking for an exciting new challenge? Join this team and transition into maintaining and enhancing the reliability of one of the world’s largest platforms. In this role, you will utilise your expertise in Golang coding to develop...SuggestedPermanent- £67k - £87k per annumEstimatedSite Reliability Engineers Position Description CGI was recognised in the Sunday Times Best Places to Work List 2025 and has been named one of the ‘World’s Best Employers’ by Forbes magazine. We offer a competitive salary, excellent pension, private healthcare, plus a...Suggested5 days/week
- £88k - £113k per annumEstimated...more people in more places. We believe that the problems we solve today unlock the opportunities of tomorrow. As a Senior Site Reliability Engineer, you’ll work to: Build and scale our internal platform offerings (compute, storage and networking services) to ensure...
- £70k - £91k per annumEstimated...London, New York, Singapore, Sydney and our newly established Engineering Hub in Lisbon. We have raised more than £500m in funding and... ...years—and a UK Best Employer for 2026. Thought Machine’s Site Reliability Engineers are the guardians of mission-critical systems for the...Full-timeImmediate startFlexible hours
£150k per annum
...Site Reliability Engineer Job Opportunity Role: Site Reliability Engineer Client: Most Elite FinTech Firm in London Compensation: Up to £150k + Bonus + Package Location: Montreal Overview An Elite FinTech Firm is looking for a highly talented DevOps...PermanentOn-siteFlexible hours£65k - £75k per annum
Senior Site Reliability Engineer Up to £75,000 plus bonus and on call allowance Milton Keynes (2 days on site a week) VIQU have partnered with a well-established B2B SaaS company who are going through a significant platform transformation. and so are hiring for a Senior...Full-timeRotating shifts2 days/week£75k per annum
...Senior Site Reliability Engineer Up to £75,000 plus bonus and on call allowance Milton Keynes (2 days on site a week) VIQU have partnered with a well-established B2B SaaS company who are going through a significant platform transformation. and so are hiring for a...Permanent2 days/week- £52k - £69k per annumEstimated...on our best-in-class platform. Our Winning Culture is the engine that drives our teams of innovators. We champion diversity of... ...join us and let’s build what’s next - together! As a Senior Site Reliability Engineer, you’ll support our production platforms, including participating...Remote jobLong-term contract
- £65k - £86k per annumEstimated...holders and more than 40 million verified users, facilitating over $1 trillion in crypto transactions. We are looking for a Site Reliability Engineer to join our Core team to encourage infrastructure best practices across our organization that would allow to securely...Full-timeApprenticeshipOn-siteRemoteFlexible hours
- ...geospatial data platform) and builds integrations that make satellite imagery easy to access and act on. We’re now looking for a Site Reliability Engineer to help us ensure our data platform is reliable, scalable, and performing at its best as we grow. What will you be doing?...Long-term contract
- £61k - £80k per annumEstimated...Serverless platform allows customers to deploy and scale their own AI models on production-grade GPU infrastructure. As a Site Reliability Engineer at Runware, you will help ensure these systems remain reliable, performant and resilient as we scale. This is a highly...Long-term contractHybrid workingRemoteWork from homeFlexible hours
£90k per annum
...Site Reliability Engineer Trading £90,000 Plus Bonus Quant Capital is urgently looking for a Site Reliability Engineer to join our high profile client. Our client has recently been voted in the fintech 100 (best financial technology businesses). They provide...Immediate startFlexible hours- £48k - £63k per annumEstimated...innovative companies, focusing on critical national services. The Role Develop and maintain resilient cloud platforms Ensure the reliability and scalability of critical national services Collaborate with development teams to implement best practices Participate in a...Full-timeShift work
- £70k - £92k per annumEstimated...of both work styles in a workplace that is intentional about belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a systems thinker. You’ll create middleware and platform guardrails...Full-timeHybrid workingOn-siteRemoteMonday to FridayFlexible hours
- £48k - £64k per annumEstimated...are building and scaling payments and card systems, ensuring our financial platform is reliable, accessible, and seamless for users. What You’ll be doing: As a Site Reliability Engineer (SRE) you will help ensure our platforms are fast, reliable, and scalable. You’ll...Hybrid workingOn-site
£95k - £110k per annum
...Without you, and the work that you do, availability, efficiency and reliability aren’t managed and everything falls apart. You’re the glue, an inspirational leader, rallying the troops to get stuck into reliability. No metric missed, all angles observed and every stone upturned...Remote jobPermanentOn-site- £52k - £68k per annumEstimatedSite Reliability Engineer Position Description We are seeking an experienced and proactive Site Reliability Engineer (SRE) to join a team supporting multiple data product and platform groups. This role is focused on improving the reliability, scalability, observability...5 days/week
- £121k - £164k per annumEstimated...Site Reliability Engineer – Fintech Quant Capital is urgently looking for a Site Reliability Engineer to join or well known Fintech50 client who produces software disrupting the wealth management market. My client is a market leading SAAS provider to financial advisory...
£45k - £55k per annum
...isn’t just a goal it’s a necessity. Stability and efficiency are key, and collaboration is at the heart of how things get done. Here, engineers work together to fix issues, improve systems, and stay at the cutting edge of technology. This hybrid role ensures you get the...PermanentHybrid working£55k - £65k per annum
...Time to enhance your scope; broaden your horizon by delving into site reliability engineering. You’ll take the skills you have picked up in software engineering and apply these to improve overall system and application performance and reliability. You’ll work on internal...PermanentHybrid working£90k - £120k per annum
...your decisions Ability to focus on what matters most, manage your time, and get things done To be a team player, ready to help engineers investigate issues or teach them new things Love to automate manual work and try new modern technology/approaches What we...Relocation packageVisa sponsorshipOn-siteRemoteFlexible hours1 day/week£7k per annum
...languages. Founded by clinicians, Heidi brings together clinicians, engineers, designers, scientists, creatives, and mathematicians, working... .... You’ll work directly on incident response, on-call, system reliability, and day-to-day operations for Heidi’s platform. We’re open...Full-timeHybrid workingOn-site- £41k - £54k per annumEstimated...down the page, but for now – let’s talk about the role and who we’re looking for… A bit about the role At Allwyn, the Site Reliability Engineer supports the reliability and performance of digital services by operating production systems, building automation, and...Full-timeFlexible hours
£50k - £70k per annum
...big step for the industry. With projects like these, it will create many new roles that will create a huge opportunity for you as an engineer to grow and progress. Alongside the large-scale projects, there is a strong focus on learning and development, providing you with budgets...PermanentHybrid working- £38k - £51k per annumEstimated...nation states, armed forces and commercial businesses can unlock digital advantage in the most demanding environments. Site Reliability Engineering is a rapidly growing concept in industry, with a remit to drive the quality, reliability and performance of essential systems...Hybrid workingOn-siteRemoteRotating shifts
£250k - £300k per annum
...mutual respect and a flat structure. This role sits in the Shared Engineering team that focuses on designing, developing, and maintaining infrastructure and tools. The team requires a Network Site Reliability Engineer (SRE) with strong network fundamentals, problem-solving...On-site- £50k - £65k per annumEstimated...At Lucera, we're seeking an experienced SRE/DevOps Engineer to join our engineering team. You'll play a crucial role in maintaining the reliability and performance of our global trading infrastructure and financial technology platforms. This role offers an opportunity to work...Full-time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- senior site reliability engineer United Kingdom
- cloud site reliability engineer United Kingdom
- site reliability engineer United Kingdom
- site reliability engineer sre
- senior site reliability engineer
- lead site reliability engineer
- cloud site reliability engineer
- site reliability engineer
- director site reliability engineering
- sre engineer

