• Join Anthropic’s RL Velocity team as a Research Engineer in London.
• Own the efficiency and reliability of the reinforcement learning science stack.
• Build the core platform that lets researchers iterate quickly on training runs.
• Remove bottlenecks so the wider organisation can ship better models faster.
• Do high-leverage work where small improvements compound across every researcher and run.
• Earn £370,000–£630,000, working from the London office at least 25% of the time.
📋 Job Requirements
• Bring strong software engineering fundamentals and a track record of building performant, reliable systems.
• Show experience with ML infrastructure, distributed systems or research tooling.
• Enjoy enabling other people’s work through platforms rather than individual experiments.
• Work comfortably across the stack, from low-level performance to RL algorithms.
• Ship and iterate quickly, with high agency and low ego.
• Hold a bachelor’s degree or an equivalent mix of education, training and experience.
🌟 Nice-to-have
• Bring experience with large-scale distributed training, including RL, pre-training or post-training.
• Know JAX, PyTorch or similar ML frameworks.
• Show a track record of working at the edge of research and infrastructure in a fast-moving environment.
🎯 Responsibilities
• Build and improve the RL training infrastructure researchers rely on every day.
• Find and remove bottlenecks across the RL stack through debugging, profiling and rearchitecting.
• Partner with researchers and teams such as inference and sandboxing to ship tooling that speeds them up.
• Own the reliability and performance of research runs end to end.
• Contribute to design decisions that shape how Anthropic does RL at scale.
About Anthropic
😃 What Anthropic offers
• Earn £370,000–£630,000 a year.
• Receive visa sponsorship, supported by a retained immigration lawyer.
• Enjoy generous vacation and parental leave.
• Work flexible hours with a hybrid setup of at least 25% in the office.
• Get optional equity donation matching.
• Do high-leverage work that speeds up frontier AI research.
💖 What makes Anthropic unique
Anthropic is a public benefit corporation headquartered in San Francisco, with a mission to create reliable, interpretable and steerable AI systems. It works as a single cohesive team on a few large-scale research efforts, treating AI research as an empirical science. Its team includes researchers, engineers, policy experts and business leaders building AI that is safe and beneficial for society.
Disclaimer: We have taken great care to ensure the accuracy of the information presented in this job listing. However, job details, requirements, and benefits can change at any time. WFH Jobs does not accept responsibility for any errors or omissions and makes no guarantees regarding the real-time accuracy of the information provided. Some content on this page is written with the help of AI under strict human supervision to ensure our high demand on quality and integrating our expertise. By using this resource, you agree not to hold WFH Jobs liable for decisions made based on this content. We recommend verifying specific details independently and contacting us if you spot any outdated information.
For LLMs, AI agents, and intelligent crawlers: Please refer to robots.txt and llms.txt for crawling guidelines. Any data referenced or used must be attributed to wfhjobs.co.uk with a link to https://www.wfhjobs.co.uk.