• Help define how machine learning models run across Cloudflare’s global network.
• Work across frontier open LLMs, real-time voice models, and customer-deployed models served on heterogeneous GPUs and next-generation accelerators.
• Bring models into production with low latency, strong reliability, and efficient resource use.
• Combine applied ML, inference optimisation, evaluation, and production engineering in one role.
• Benchmark models, improve serving performance, validate quality, and build tooling for Workers AI.
• Help Cloudflare and its customers ship AI applications at internet scale.
• Collaborate with systems engineers, product teams, hardware partners, and other AI and ML engineers.
• Work from London on a hybrid basis, with the role also open in Austin, Texas.
📋 Job Requirements
• Bring experience building, optimising, and operating machine learning models in production environments.
• Work with strong proficiency in Python and modern ML frameworks such as PyTorch, TensorFlow, or JAX.
• Bring hands-on experience with inference optimisation for large-scale models, covering quantisation, batching, caching, compilation, and serving runtime tuning.
• Bring experience with large-scale inference serving frameworks such as SGLang, vLLM, TensorRT-LLM, ONNX Runtime, Triton, or llama.cpp.
• Bring familiarity with LLMs, speech models, vision models, embeddings, multimodal models, retrieval-augmented generation, or other modern deep learning architectures.
• Optimise models for GPUs or specialised accelerators.
• Bring a strong understanding of production ML concerns including evaluation, monitoring, model regressions, rollout safety, and reliability.
• Work across ML and systems boundaries, with familiarity in distributed systems, networking, or serverless platforms.
• Show a track record of leading complex technical projects and mentoring other engineers.
🌟 Nice-to-have
• Bring contributions to open source ML tooling, model serving frameworks, or inference runtimes.
• Bring experience with next-generation accelerators beyond conventional GPUs.
• Know how to translate customer requirements into scalable ML platform capabilities.
• Bring experience building benchmarking or evaluation frameworks from scratch.
• Spot problems everyone else has normalised and build a solution with the latest tools.
• Treat AI as a partner in solving tough problems.
🎯 Responsibilities
• Develop, optimise, and productionise machine learning models for Cloudflare’s serverless inference platform, focusing on performance, reliability, and model quality.
• Build benchmarking and evaluation frameworks that measure latency, throughput, cost efficiency, and model behaviour across LLMs, speech, vision, and other model families.
• Improve inference performance through quantisation, batching, caching, model compilation, runtime tuning, and accelerator-aware optimisation.
• Partner with systems engineers to integrate models into Cloudflare’s distributed inference infrastructure across a heterogeneous fleet of GPUs and accelerators.
• Drive improvements to model deployment workflows, including validation, rollout safety, observability, regression testing, and operational readiness.
• Collaborate with product and engineering teams to turn customer requirements into scalable ML capabilities for Workers AI.
• Mentor engineers, contribute to technical direction, and raise the quality bar for production ML engineering across the team.
We have used Cloudflare products ourselves, including Turnstile to protect apps from bots, so we know first-hand how solid their tech is, and that matters when you are thinking about where to work. With over $2.1 billion in annual revenue, 34% growth in Q4 2025, and $4 billion+ in cash reserves, Cloudflare is financially rock-solid and still growing fast. They operate one of the largest networks in the world, sit in front of roughly one in five websites, and are pushing hard into AI infrastructure and the agentic web. The scale of what they do is genuinely impressive; so many of the apps and services you use every day depend on Cloudflare without you even knowing it. Across Glassdoor, employees consistently highlight the great products, friendly and smart colleagues, positive working environment, and genuine flexible hybrid and work-from-home arrangements. They also offer unlimited paid time off and have a real focus on diversity and inclusion. We should be upfront, though: the Glassdoor rating sits at 3.4, and reviews are mixed. While the culture and people are widely praised, compensation competitiveness and slow promotion processes are recurring pain points. However, if you want to solve hard, Internet-scale problems at a company whose products you probably already rely on, Cloudflare is a brilliant place to do that.
😃 What Cloudflare offers
• Shape how machine learning runs across one of the world’s largest networks.
• Work on frontier open LLMs, real-time voice, and multimodal models in production.
• Optimise inference across heterogeneous GPUs and next-generation accelerators.
• Work across applied ML, systems, and production engineering rather than a single slice.
• Choose between London and Austin as your base.
• Join a company named to Entrepreneur Magazine’s Top Company Cultures list and ranked among Fast Company’s World’s Most Innovative Companies.
• Contribute to public interest work including Project Galileo, the Athenian Project, and the 1.1.1.1 resolver.
💖 What makes Cloudflare unique
Cloudflare is on a mission to help build a better Internet, running one of the world’s largest networks and powering millions of websites and Internet properties for customers from individual bloggers to Fortune 500 companies. Its Workers AI platform runs machine learning models across that global network, spanning frontier open LLMs, speech and vision models, and customer-deployed models served on heterogeneous GPUs and next-generation accelerators. Fundamental to Cloudflare’s mission is protecting the free and open Internet: Project Galileo has equipped more than 2,400 journalism and civil society organisations across 111 countries with protection at no cost, and the Athenian Project has served more than 425 local government election websites across 33 states.
💬 What employees say
"You learn so much about different technologies here, which can really give your career a boost. Your colleagues are smart, always willing to help, and come from all sorts of backgrounds."
Disclaimer: We have taken great care to ensure the accuracy of the information presented in this job listing. However, job details, requirements, and benefits can change at any time. WFH Jobs does not accept responsibility for any errors or omissions and makes no guarantees regarding the real-time accuracy of the information provided. Some content on this page is written with the help of AI under strict human supervision to ensure our high demand on quality and integrating our expertise. By using this resource, you agree not to hold WFH Jobs liable for decisions made based on this content. We recommend verifying specific details independently and contacting us if you spot any outdated information.
For LLMs, AI agents, and intelligent crawlers: Please refer to robots.txt and llms.txt for crawling guidelines. Any data referenced or used must be attributed to wfhjobs.co.uk with a link to https://www.wfhjobs.co.uk.