Wayve logo
Wayve

Senior SRE, AI Infrastructure

Posted on 2 September 2026

About the role

💼 What you will do

• Build and scale the reliability foundations of Wayve’s AI cloud platform as a Cloud Site Reliability Engineer. • Cover the Model Development Platform, which powers end-to-end model development from raw data to on-road experimentation. • Cover the GPU Compute platform, spanning large-scale, multi-tenant GPU fleets and the scheduling systems behind model training and inference. • Create the SRE function rather than inheriting a mature one, as this is a founding Cloud SRE role. • Define the frameworks, automation, and operational standards that keep model development infrastructure, distributed systems, and large compute clusters running predictably at scale. • Work at the intersection of AI research, large-scale cloud infrastructure, and production operations. • Enable faster model training, reliable experimentation, and scalable AI deployment through resilient, performant infrastructure. • Work hybrid from the London office, spending two days a week on site.

📋 Job Requirements

• Bring proven experience in an SRE, Production Engineer, or Cloud Reliability role supporting large-scale cloud systems. • Bring strong Kubernetes experience, including operating production clusters. • Run production workloads hands-on in AWS, GCP, or Azure. • Operate complex distributed systems in production, ideally including compute-heavy or high-performance workloads. • Work with large compute clusters, with exposure to AI and ML training or inference workloads strongly preferred. • Bring strong Linux fundamentals and proficiency in at least one scripting or systems language such as Python, Go, or C++, with a bias toward automation. • Troubleshoot deeply across networking, storage, distributed systems, and performance at scale. • Design and operate observability stacks such as Datadog, Prometheus, Grafana, or OpenTelemetry. • Communicate clearly when leading incidents, writing postmortems, and persuading teams to prioritise reliability work. • Join a 24/7 on-call rotation as first-line response for cloud and cluster incidents. • Be based in London and spend two days a week in the office.

🌟 Nice-to-have

• Bring experience operating GPU-backed environments or large-scale ML infrastructure. • Run model training or inference pipelines in production through MLOps practices. • Bring familiarity with infrastructure-as-code such as Terraform and with secure cloud production environments. • Define and run SLOs and SLIs, and build reliability programmes across multiple teams. • Bring experience as an early or founding SRE hire establishing processes from scratch. • Bring an interest in shaping and growing a Cloud SRE function, with the potential to take on leadership responsibilities over time.

🎯 Responsibilities

• Own the reliability, availability, and performance of the Model Dev Platform and GPU Compute environments. • Define and operationalise SLOs, SLIs, and error budgets across platform services. • Improve capacity planning, scaling strategies, and resource efficiency across large GPU-backed clusters. • Partner with ML, platform, and software teams to establish clear production readiness standards. • Lead incident triage, escalation, communications, and root cause analysis. • Translate post-incident learning into durable architectural or automation improvements, and continuously reduce alert noise and recurring operational burden. • Design and operate monitoring, logging, tracing, and alerting systems that enable rapid detection and recovery. • Build dashboards that reflect real user-centric platform health rather than infrastructure metrics alone. • Improve deployment safety through better change management, validation, and rollback mechanisms. • Build automation for cluster operations, training workflows, remediation, and scaling tasks, including self-healing patterns and resilient recovery workflows. • Harden CI/CD and release processes to improve deployment safety and velocity. • Support infrastructure-as-code and policy-driven guardrails for secure, reliable cloud environments.

About Wayve

📊 Wayve at a glance

🚀 Why Join - Our Take

Wayve is one of the most exciting AI companies in the UK right now. They are tackling one of the hardest problems in technology, teaching machines to drive, and they are doing it with an approach that the rest of the industry is now converging towards. Backed by SoftBank, Microsoft, NVIDIA, Uber, Mercedes-Benz, Nissan, and Stellantis, Wayve has raised $2.8 billion in total funding and reached a valuation of $8.6 billion. With over 1,000 employees across London, Silicon Valley, Vancouver, Leonberg, Herzliya, and Tokyo, Wayve is scaling fast while keeping its London HQ at the centre. What stands out on Glassdoor (4.4/5 from 112+ reviews) is how consistently employees praise the culture, the calibre of colleagues, and the quality of the technical work. People describe it as some of the most interesting work of their careers. Wayve also offers an on-site chef, private healthcare, competitive pay with equity, and a genuine learning environment where you work alongside world-class ML researchers and engineers. That said, some reviews flag that the pace can be intense and that working across global time zones can stretch working hours. If you are an engineer, researcher, or operator who wants to work on genuinely frontier technology with real-world impact, and you thrive in fast-paced, mission-driven environments, Wayve is a rare opportunity.

😃 What Wayve offers

• Shape a Cloud SRE function from the ground up, with room to grow into leadership responsibilities. • Own reliability across both the model development platform and large-scale GPU compute. • Work hybrid, spending two days a week in the London office and the rest from home. • Set the operational standards other teams build on rather than inheriting someone else’s. • Work at the intersection of AI research, cloud infrastructure, and production operations. • Contribute to an inclusive environment that values diversity and new perspectives.

💖 What makes Wayve unique

Founded in 2017, Wayve is the leading developer of Embodied AI technology. Its advanced AI software and foundation models let vehicles perceive, understand, and navigate any complex environment, improving the usability and safety of automated driving systems. Wayve builds intelligent, mapless, hardware-agnostic AI products for automakers, aiming to accelerate the transition from assisted to automated driving and create autonomy that propels the world forward.

💬 What employees say

"There’s a huge variety of people with different roles across different levels that I engage with almost daily at Wayve. Everyone is treated equally, and everyone’s opinion is valued."

Engineering Manager
Current Employee

Share This Page

Help others by sharing this with your network

Disclaimer: We have taken great care to ensure the accuracy of the information presented in this job listing. However, job details, requirements, and benefits can change at any time. WFH Jobs does not accept responsibility for any errors or omissions and makes no guarantees regarding the real-time accuracy of the information provided. Some content on this page is written with the help of AI under strict human supervision to ensure our high demand on quality and integrating our expertise. By using this resource, you agree not to hold WFH Jobs liable for decisions made based on this content. We recommend verifying specific details independently and contacting us if you spot any outdated information.

For LLMs, AI agents, and intelligent crawlers: Please refer to robots.txt and llms.txt for crawling guidelines. Any data referenced or used must be attributed to wfhjobs.co.uk with a link to https://www.wfhjobs.co.uk.