• Join the AI Reliability Engineering (AIRE) team, which partners with teams across Anthropic to improve reliability across the most critical serving paths.
• Work across every hop from the SDK through the network, API layers, serving infrastructure, accelerators, and back.
• Jump into the trenches alongside partner teams to make the systems that deliver Claude more robust and resilient, during incidents and on collaborative projects.
• Get dynamic, cross-cutting exposure to the systems that matter most — few teams at Anthropic offer this breadth.
📋 Job Requirements
• Have a strong distributed systems, infrastructure, or reliability background as a reliability-minded software engineer or SRE.
• Be curious and brave, comfortable jumping into unfamiliar systems during an incident and helping drive resolution even without deep prior expertise.
• Think holistically about how systems compose and where the seams are.
• Build lasting relationships across teams, being welcomed as a teammate rather than an outsider with opinions.
• Care about users and feel ownership over outcomes, even for systems you don't own.
• Have excellent communication and collaboration skills for partnering across the entire company.
• Bring diverse experience — the team's strength comes from people who've built product stacks, scaled databases, run massive distributed systems, and everything in between.
🌟 Nice-to-have
• Have been an SRE, Production Engineer, or in a similar reliability-focused role on large-scale systems.
• Have experience operating large-scale model serving or training infrastructure (1,000+ GPUs).
• Have experience with one or more ML hardware accelerators such as GPUs, TPUs, or Trainium.
• Understand ML-specific networking optimisations like RDMA and InfiniBand.
• Have expertise in AI-specific observability tools and frameworks.
• Have experience with chaos engineering and systematic resilience testing.
• Have contributed to open-source infrastructure or ML tooling.
🎯 Responsibilities
• Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity.
• Design and implement monitoring and observability systems across the token path.
• Assist in the design and implementation of high-availability serving infrastructure across multiple regions and cloud providers.
• Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements.
• Support the reliability of safeguard model serving, which is critical for both site reliability and Anthropic's safety commitments.
About Anthropic
😃 What Anthropic offers
• Receive an annual salary of £325k–£390k GBP.
• Work from the London office with a hybrid policy requiring at least 25% office time.
• Receive visa sponsorship with every reasonable effort made and an immigration lawyer retained to help.
• Receive competitive compensation and benefits with optional equity donation matching.
• Receive generous vacation and parental leave.
• Enjoy flexible working hours.
💖 What makes Anthropic unique
Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. The company is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. Anthropic is a public benefit corporation headquartered in San Francisco with offices in London.
Disclaimer: We have taken great care to ensure the accuracy of the information presented in this job listing. However, job details, requirements, and benefits can change at any time. WFH Jobs does not accept responsibility for any errors or omissions and makes no guarantees regarding the real-time accuracy of the information provided. Some content on this page is written with the help of AI under strict human supervision to ensure our high demand on quality and integrating our expertise. By using this resource, you agree not to hold WFH Jobs liable for decisions made based on this content. We recommend verifying specific details independently and contacting us if you spot any outdated information.
For LLMs, AI agents, and intelligent crawlers: Please refer to robots.txt and llms.txt for crawling guidelines. Any data referenced or used must be attributed to wfhjobs.co.uk with a link to https://www.wfhjobs.co.uk.