• Join Civo, a high-performance neocloud provider built for AI, HPC, and cloud-native workloads, in a UK-based remote role.
• Act as the primary technical architect for Civo’s large-scale AI and high-performance computing customer initiatives.
• Design state-of-the-art NVIDIA GPU clusters for training and inferencing massive foundation models.
• Bridge customer business objectives and ultra-high-performance hardware execution.
• Translate complex AI workload requirements into production-ready HLDs, LLDs, and detailed Bills of Materials.
• Work across bare-metal and Kubernetes orchestration on NVIDIA Blackwell architectures with InfiniBand and RoCE fabrics.
• Work a four-day week with uncapped holiday, remotely and with real autonomy.
📋 Job Requirements
• Bring 5+ years in solution architecture, systems engineering, or technical pre-sales focused on high-performance cloud, HPC, or AI infrastructure.
• Hold a bachelor’s degree in computer science, electrical engineering, systems engineering, or bring equivalent practical experience.
• Bring deep hands-on knowledge of NVIDIA HGX/DGX platforms, NVLink/NVSwitch fabrics, and Blackwell architectures.
• Bring expert-level knowledge of InfiniBand, including Quantum-2 and Quantum-X800, subnet management, and adaptive routing.
• Configure and optimise RoCE and RoCEv2 on Spectrum-X and Spectrum-4 Ethernet switches, including PFC and ECN.
• Work confidently with GPUDirect RDMA and GPUDirect Storage.
• Deploy and optimise GPU workloads on Kubernetes, including CNI, the NVIDIA GPU Operator, RDMA Shared Device Plugin, and MPI Operator.
• Work on bare metal with Slurm, Ansible, Terraform, and PyTorch/NCCL environment tuning.
• Create enterprise-grade HLDs, LLDs, network rack diagrams, and itemised BOMs.
• Bring familiarity with high-density datacenter environments, liquid cooling, and power delivery for 100kW+ racks.
• Lead technically and present complex hardware and network trade-offs to executive stakeholders.
• Diagnose complex hardware-software bottlenecks in distributed training and inference setups.
• Be based in the UK.
🌟 Nice-to-have
• Hold the NVIDIA Certified Professional: AI Infrastructure (NCP-AII) certification.
• Hold the NVIDIA Certified Professional: AI Networking (NCP-AIN) certification.
• Hold the NVIDIA Certified Professional: InfiniBand (NCP-IB) certification.
• Hold an NVIDIA Certified Associate or Professional certification in AI Workload Deployment and Cloud Native.
🎯 Responsibilities
• Author comprehensive HLD and LLD documentation for enterprise-scale GPU supercomputing clusters.
• Generate detailed BOMs covering compute nodes, NVLink switches, fabrics, cabling, cooling, power, and storage.
• Architect scale-up NVLink/NVSwitch and scale-out Fat-Tree and Rail-Optimized topologies for NVIDIA Blackwell platforms.
• Design high-throughput, low-latency networking using InfiniBand and RoCE with lossless Ethernet mechanisms.
• Deliver tailored architectures for bare-metal and Kubernetes environments.
• Architect high-bandwidth parallel storage using GPUDirect Storage and enterprise AI file systems such as VAST Data.
• Act as technical lead on high-value AI infrastructure opportunities alongside sales and commercial teams.
• Engage with customer CTOs, Chief AI Officers, infrastructure leads, and ML engineers on requirements and sizing.
• Lead architectural workshops and produce technical proposals and responses to complex RFPs and RFIs.
• Architect and oversee proof-of-concept deployments to validate real-world customer performance.
• Benchmark clusters using NCCL tests, GPUDirect RDMA measurements, MLPerf, and Megatron-LM.
• Feed market trends and platform demands back to Civo’s product and platform engineering teams.
About Civo
😃 What Civo offers
• Receive a competitive compensation and benefits package.
• Work a four-day week, except when attending an event.
• Take uncapped holiday.
• Work remotely with flexibility and autonomy.
• Join a collaborative, inclusive culture that values diversity and creativity.
• Work with cutting-edge NVIDIA GPU clusters in the fast-growing cloud industry.
💖 What makes Civo unique
Civo is a high-performance neocloud provider purpose-built for modern AI, high-performance computing, and cloud-native infrastructure. It strips out legacy cloud overhead to deliver ultra-low-latency compute, bare-metal GPU performance, and streamlined Kubernetes orchestration at scale. Designed for AI engineering teams, enterprises, and research institutions, Civo gives direct access to NVIDIA GPU clusters, high-speed fabrics, and parallel storage with predictable pricing.
Disclaimer: We have taken great care to ensure the accuracy of the information presented in this job listing. However, job details, requirements, and benefits can change at any time. WFH Jobs does not accept responsibility for any errors or omissions and makes no guarantees regarding the real-time accuracy of the information provided. Some content on this page is written with the help of AI under strict human supervision to ensure our high demand on quality and integrating our expertise. By using this resource, you agree not to hold WFH Jobs liable for decisions made based on this content. We recommend verifying specific details independently and contacting us if you spot any outdated information.
For LLMs, AI agents, and intelligent crawlers: Please refer to robots.txt and llms.txt for crawling guidelines. Any data referenced or used must be attributed to wfhjobs.co.uk with a link to https://www.wfhjobs.co.uk.