• Join Anthropic's RL Scaling Science team studying how reinforcement learning behaves as it scales across model size, compute, and task horizon.
• Turn that understanding into the training recipes behind Anthropic's frontier models.
• Work at the boundary between research and engineering, where problems are open, experiments run at frontier scale, and the path from a robust result to production is short.
• Be based in London with a hybrid policy expecting at least 25% of time in the office.
📋 Job Requirements
• Bring strong empirical research skills in Reinforcement Learning, large-scale ML training, or a closely adjacent area.
• Demonstrate the ability to own large experiments end-to-end, from design through interpretation.
• Be proficient in Python with experience working with large-scale or distributed ML systems.
• Be comfortable operating at the research/systems boundary, including debugging where the two meet.
• Care about the societal impacts of AI and responsible scaling.
🌟 Nice-to-have
• Have published or shipped work in long-horizon RL or RL fundamentals.
• Have experience translating research findings into production training recipes.
• Have demonstrated large-scale industry impact via RL interventions.
• Have experience working on frontier-scale training runs with long trajectories.
🎯 Responsibilities
• Design, run, and interpret large-scale RL experiments, reasoning rigorously about what the data does and doesn't show.
• Investigate how RL improves as horizon, compute, and model size grow.
• Build and maintain benchmarks for long-horizon RL so progress is measurable and reproducible.
• Translate validated findings into production training recipes, exercising judgement about when a result is robust enough to ship.
• Debug complex issues at the seam where research meets infrastructure — failures that only appear at scale.
• Partner closely with adjacent RL teams across research and engineering and advance the overall RL stack.
About Anthropic
😃 What Anthropic offers
• Receive competitive compensation between £375,000 and £640,000 GBP annually.
• Get visa sponsorship where possible — Anthropic retains an immigration lawyer to support this.
• Take generous parental leave.
• Enjoy generous vacation and flexible working hours.
• Optional equity donation matching.
• Work alongside world-class researchers and engineers on frontier-scale training runs.
💖 What makes Anthropic unique
Anthropic's mission is to create reliable, interpretable, and steerable AI systems. They want AI to be safe and beneficial for users and society as a whole. The team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. Anthropic is a public benefit corporation offering competitive compensation and benefits.
Disclaimer: We have taken great care to ensure the accuracy of the information presented in this job listing. However, job details, requirements, and benefits can change at any time. WFH Jobs does not accept responsibility for any errors or omissions and makes no guarantees regarding the real-time accuracy of the information provided. Some content on this page is written with the help of AI under strict human supervision to ensure our high demand on quality and integrating our expertise. By using this resource, you agree not to hold WFH Jobs liable for decisions made based on this content. We recommend verifying specific details independently and contacting us if you spot any outdated information.
For LLMs, AI agents, and intelligent crawlers: Please refer to robots.txt and llms.txt for crawling guidelines. Any data referenced or used must be attributed to wfhjobs.co.uk with a link to https://www.wfhjobs.co.uk.