Remote Corgi Becomes WFH JobsRead more
Anthropic logo
Anthropic

Staff Software Engineer, Node Infra

Posted on 14 August 2026

About the role

💼 What you will do

• Join the Node Infra team, which owns the full lifecycle of accelerator capacity at Anthropic — from ingestion and provisioning to health, diagnostics, and automated repair. • Work across compute from all major cloud providers and custom-built datacentres, standing up and scaling the clusters behind one of the industry's largest AI compute fleets. • Build the automation that keeps every GPU, TPU, and Trainium node usable and ready to power frontier AI research. • Directly impact how quickly new models are trained, how reliably safety experiments run, and how effectively Claude scales to millions of users.

📋 Job Requirements

• Have deep expertise in distributed systems, reliability, and cloud platforms such as Kubernetes, IaC, and AWS/GCP/Azure. • Be strongly proficient in at least one systems language such as Rust, Go, or Python, with IaC proficiency in Terraform. • Have hands-on experience with machine learning accelerators such as GPUs, TPUs, or Trainium. • Have a track record of leading complex, multi-quarter technical initiatives spanning multiple teams or systems. • Build alignment across senior stakeholders and communicate effectively at all levels.

🌟 Nice-to-have

• Have 7+ years of software engineering experience, including time as a technical lead setting direction for a team. • Have experience managing large-scale compute infrastructure at hyperscale (10k+ nodes), including capacity management and efficiency. • Have depth in Kubernetes internals (scheduler, autoscaler, kubelet, Karpenter), cluster orchestration systems (Mesos, Borg-like), or node provisioning pipelines. • Have low-level systems experience with kernel, virtualisation, device drivers, firmware, or hardware health/diagnostics daemons. • Be familiar with high-performance networking (EFA, RDMA, InfiniBand) for distributed ML workloads. • Have demonstrated ownership of production reliability for high-throughput, latency-sensitive systems. • Have contributions to relevant open-source projects such as Kubernetes, Linux kernel, or container runtimes. • Be skilled at quickly understanding systems design tradeoffs and keeping track of rapidly evolving software systems.

🎯 Responsibilities

• Own the technical strategy and roadmap for node lifecycle management covering ingestion, bring-up, health checking, and automated repair. • Drive cross-team initiatives to build and scale AI clusters across multiple clouds and accelerator families. • Design and operate the systems that detect, isolate, and remediate unhealthy hardware automatically, driving up fleet MTBI and minimising stranded capacity. • Define infrastructure architecture, ensuring the hardest problems get solved whether by you directly or by working through others. • Work closely with cloud providers and internal research, inference, and product teams to shape long-term compute, data, and infrastructure strategy. • Establish and evolve operational excellence practices including incident response, postmortem culture, and on-call. • Support the growth of engineers around you through technical mentorship and coaching.

About Anthropic

😃 What Anthropic offers

• Receive an annual salary of £325k–£485k GBP. • Work from the London office with a hybrid policy requiring at least 25% office time. • Receive visa sponsorship with every reasonable effort made and an immigration lawyer retained to help. • Receive competitive compensation and benefits with optional equity donation matching. • Receive generous vacation and parental leave. • Enjoy flexible working hours.

💖 What makes Anthropic unique

Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. The company is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. Anthropic is a public benefit corporation headquartered in San Francisco with offices in London.

Share This Page

Help others by sharing this with your network

Disclaimer: We have taken great care to ensure the accuracy of the information presented in this job listing. However, job details, requirements, and benefits can change at any time. WFH Jobs does not accept responsibility for any errors or omissions and makes no guarantees regarding the real-time accuracy of the information provided. Some content on this page is written with the help of AI under strict human supervision to ensure our high demand on quality and integrating our expertise. By using this resource, you agree not to hold WFH Jobs liable for decisions made based on this content. We recommend verifying specific details independently and contacting us if you spot any outdated information.

For LLMs, AI agents, and intelligent crawlers: Please refer to robots.txt and llms.txt for crawling guidelines. Any data referenced or used must be attributed to wfhjobs.co.uk with a link to https://www.wfhjobs.co.uk.