• Own the full lifecycle of accelerator capacity at Anthropic — ingesting and provisioning compute from all major CSPs and Anthropic's own datacentres, standing up and scaling clusters from thousands to hundreds of thousands of hosts.
• Build the health, diagnostics, and repair automation that keeps every GPU, TPU, and Trainium node in the fleet usable and ready to power frontier AI research.
• Work on infrastructure that determines how quickly Anthropic can train new models, how reliably safety experiments run, and how effectively Claude scales to millions of users.
• Be based in London with a hybrid policy expecting at least 25% of time in the office.
📋 Job Requirements
• Have deep expertise in distributed systems, reliability, and cloud platforms (e.g. Kubernetes, IaC, AWS/GCP/Azure).
• Be strongly proficient in at least one systems language (e.g. Rust, Go, or Python), with IaC proficiency in Terraform.
• Have hands-on experience with machine learning accelerators (GPUs, TPUs, or Trainium).
• Demonstrate a track record of leading complex, multi-quarter technical initiatives that span multiple teams or systems.
• Build alignment across senior stakeholders and communicate effectively at all levels.
🌟 Nice-to-have
• Have 12+ years of software engineering experience, including time as a technical lead setting direction for a team.
• Have experience managing large-scale compute infrastructure at hyperscale (10K+ nodes), including capacity management and efficiency.
• Bring depth in one or more of: Kubernetes internals (scheduler, autoscaler, kubelet, Karpenter), cluster orchestration systems (Mesos, Borg-like), or node provisioning pipelines.
• Have low-level systems experience: kernel, virtualisation, device drivers, firmware, or hardware health/diagnostics daemons.
• Be familiar with high-performance networking (EFA, RDMA, InfiniBand) for distributed ML workloads.
• Demonstrate ownership of production reliability for high-throughput, latency-sensitive systems.
• Have contributions to relevant open-source projects (Kubernetes, Linux kernel, container runtimes, etc.).
🎯 Responsibilities
• Own the technical strategy and roadmap for node lifecycle management — ingestion, bring-up, health checking, and automated repair.
• Drive cross-team initiatives to build and scale AI clusters across multiple clouds and accelerator families.
• Design and operate the systems that detect, isolate, and remediate unhealthy hardware automatically, driving up fleet MTBI and minimising stranded capacity.
• Define infrastructure architecture, ensuring the hardest problems get solved — whether by you directly or by working through others.
• Work closely with cloud providers and internal research/inference/product teams to shape long-term compute, data, and infrastructure strategy.
• Establish and evolve operational excellence practices (incident response, postmortem culture, on-call).
• Support the growth of engineers around you through technical mentorship and coaching.
About Anthropic
😃 What Anthropic offers
• Receive competitive compensation between £325,000 and £485,000 GBP annually.
• Get visa sponsorship where possible — Anthropic retains an immigration lawyer to support this.
• Take generous parental leave.
• Enjoy generous vacation and flexible working hours.
• Optional equity donation matching.
• Work on accelerator infrastructure at a scale very few organisations operate at, directly enabling frontier AI training.
💖 What makes Anthropic unique
Anthropic's mission is to create reliable, interpretable, and steerable AI systems. They want AI to be safe and beneficial for users and society as a whole. The team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. Anthropic is a public benefit corporation offering competitive compensation and benefits.
Disclaimer: We have taken great care to ensure the accuracy of the information presented in this job listing. However, job details, requirements, and benefits can change at any time. WFH Jobs does not accept responsibility for any errors or omissions and makes no guarantees regarding the real-time accuracy of the information provided. Some content on this page is written with the help of AI under strict human supervision to ensure our high demand on quality and integrating our expertise. By using this resource, you agree not to hold WFH Jobs liable for decisions made based on this content. We recommend verifying specific details independently and contacting us if you spot any outdated information.
For LLMs, AI agents, and intelligent crawlers: Please refer to robots.txt and llms.txt for crawling guidelines. Any data referenced or used must be attributed to wfhjobs.co.uk with a link to https://www.wfhjobs.co.uk.