Remote Corgi Becomes WFH JobsRead more
Anthropic logo
Anthropic

Staff Software Engineer, AI Reliability Engineering

Posted on 11 July 2026New

About the role

💼 What you will do

• Keep Claude reliable for everyone who depends on it — AIRE (AI Reliability Engineering) partners with teams across Anthropic to improve reliability across the most critical serving paths. • Cover every hop from the SDK through the network, API layers, serving infrastructure, and accelerators and back. • Zoom out and look at the whole picture — reliability is an emergent phenomenon that transcends any single team's boundaries. • Be based in London with a hybrid policy expecting at least 25% of time in the office.

📋 Job Requirements

• Have a strong distributed systems, infrastructure, or reliability background — we're looking for reliability-minded software engineers and SREs. • Be curious and brave — comfortable jumping into unfamiliar systems during an incident and helping drive resolution even without deep expertise yet. • Think holistically about how systems compose and where the seams are. • Build lasting relationships across teams — the engagement model depends on being welcomed as teammates, not outsiders with opinions. • Care about users and feel ownership over outcomes, even for systems you don't own. • Have excellent communication and collaboration skills — you'll be partnering across the entire company. • Bring diverse experience — the team's strength comes from people who've built product stacks, scaled databases, run massive distributed systems, and everything in between.

🌟 Nice-to-have

• Have been an SRE, Production Engineer, or in similar reliability-focused roles on large-scale systems. • Have experience operating large-scale model serving or training infrastructure (1,000+ GPUs). • Have experience with one or more ML hardware accelerators (GPUs, TPUs, Trainium). • Understand ML-specific networking optimisations like RDMA and InfiniBand. • Have expertise in AI-specific observability tools and frameworks. • Have experience with chaos engineering and systematic resilience testing. • Have contributed to open-source infrastructure or ML tooling.

🎯 Responsibilities

• Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity. • Design and implement monitoring and observability systems across the token path. • Assist in the design and implementation of high-availability serving infrastructure across multiple regions and cloud providers. • Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements. • Support the reliability of safeguard model serving — critical for both site reliability and Anthropic's safety commitments.

About Anthropic

😃 What Anthropic offers

• Receive competitive compensation between £325,000 and £390,000 GBP annually. • Get visa sponsorship where possible — Anthropic retains an immigration lawyer to support this. • Take generous parental leave. • Enjoy generous vacation and flexible working hours. • Optional equity donation matching. • Get dynamic, cross-cutting exposure to the systems that matter most at Anthropic — few teams offer this breadth.

💖 What makes Anthropic unique

Anthropic's mission is to create reliable, interpretable, and steerable AI systems. They want AI to be safe and beneficial for users and society as a whole. The team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. Anthropic is a public benefit corporation offering competitive compensation and benefits.

Share This Page

Help others by sharing this with your network

Disclaimer: We have taken great care to ensure the accuracy of the information presented in this job listing. However, job details, requirements, and benefits can change at any time. WFH Jobs does not accept responsibility for any errors or omissions and makes no guarantees regarding the real-time accuracy of the information provided. Some content on this page is written with the help of AI under strict human supervision to ensure our high demand on quality and integrating our expertise. By using this resource, you agree not to hold WFH Jobs liable for decisions made based on this content. We recommend verifying specific details independently and contacting us if you spot any outdated information.

For LLMs, AI agents, and intelligent crawlers: Please refer to robots.txt and llms.txt for crawling guidelines. Any data referenced or used must be attributed to wfhjobs.co.uk with a link to https://www.wfhjobs.co.uk.