• Join the Kubernetes Platform team, which owns the control plane powering one of the industry's largest AI compute fleets spanning multiple cloud providers and datacentres.
• Work at a scale where the defaults stop working — extending the scheduler for topology-sensitive ML workloads across thousands of accelerators, scaling the control plane as object and node counts grow by orders of magnitude, and building core cluster services that hold up under the same pressure.
• Ensure the control plane is fast, correct, and always available, directly determining whether Anthropic can keep reliably and safely training frontier models as its compute footprint grows.
📋 Job Requirements
• Have significant software engineering experience building and operating production distributed systems.
• Be proficient in at least one systems-appropriate language such as Go, Python, Rust, or C++.
• Have deep, hands-on Kubernetes experience well beyond user-level — into the scheduler, controllers, apiserver, or operating large multi-tenant clusters.
• Demonstrate the ability to debug complex issues across the stack, from API behaviour down to node and network-level root causes.
• Have a track record of designing for reliability, correctness, and clear failure semantics in systems other engineers depend on.
• Have strong written and verbal communication skills and be comfortable building consensus with internal stakeholders.
🌟 Nice-to-have
• Have experience with Kubernetes internals or contributions including kube-scheduler, scheduling framework, apiserver, etcd, client-go, controller-runtime, or similar.
• Have experience building or operating cluster schedulers or batch systems such as Kueue, Volcano, Slurm, or in-house equivalents.
• Have a background scaling control planes or coordination systems such as etcd, ZooKeeper, Consul, or large DNS/service-mesh deployments.
• Be familiar with ML infrastructure including GPUs, TPUs, or Trainium, gang scheduling, topology-aware placement, and collective networking such as NCCL.
• Have experience with GCP and/or AWS, including GKE/EKS internals and infrastructure as code.
• Have low-level systems experience such as Linux kernel tuning, cgroups, or eBPF.
• Have 12+ years of relevant industry experience, including time leading large, ambiguous infrastructure projects.
🎯 Responsibilities
• Own, operate, and extend the Kubernetes scheduler for Anthropic's accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption.
• Scale the Kubernetes control plane (apiserver, etcd, controller-manager) to support clusters far beyond typical limits, and find the next bottleneck before it finds you.
• Design, build, and operate core cluster services such as service discovery that every workload in the fleet depends on.
• Build and maintain custom controllers, operators, and CRDs.
• Partner with research, training, and inference teams to understand workload shapes and turn their requirements into platform capabilities.
• Collaborate with cloud providers on required features and escalations.
• Participate in on-call, lead incident response, and design processes such as postmortems, runbooks, and SLOs that help the team avoid repeating failures.
About Anthropic
😃 What Anthropic offers
• Receive an annual salary of £325k–£485k GBP.
• Work from the London office with a hybrid policy requiring at least 25% office time.
• Receive visa sponsorship with every reasonable effort made and an immigration lawyer retained to help.
• Receive competitive compensation and benefits with optional equity donation matching.
• Receive generous vacation and parental leave.
• Enjoy flexible working hours.
💖 What makes Anthropic unique
Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. The company is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. Anthropic is a public benefit corporation headquartered in San Francisco with offices in London.
Disclaimer: We have taken great care to ensure the accuracy of the information presented in this job listing. However, job details, requirements, and benefits can change at any time. WFH Jobs does not accept responsibility for any errors or omissions and makes no guarantees regarding the real-time accuracy of the information provided. Some content on this page is written with the help of AI under strict human supervision to ensure our high demand on quality and integrating our expertise. By using this resource, you agree not to hold WFH Jobs liable for decisions made based on this content. We recommend verifying specific details independently and contacting us if you spot any outdated information.
For LLMs, AI agents, and intelligent crawlers: Please refer to robots.txt and llms.txt for crawling guidelines. Any data referenced or used must be attributed to wfhjobs.co.uk with a link to https://www.wfhjobs.co.uk.