• Own the data foundations that power Canva's multimodal agent research.
• Build the pipelines, datasets, and tooling that turn ambitious research ideas into trainable reality.
• Take responsibility for the full data lifecycle, from collection and curation through preprocessing, quality assurance, and delivery into training pipelines.
• Work closely with research scientists to understand what data is needed, then design and build the systems that deliver it reliably and at scale.
• Enjoy significant autonomy over how data problems get solved while aligning with the team on which problems matter most.
• Join a cutting-edge research team exploring multimodal agentic architectures, pre- and post-training, and design agents.
• Work from the London campus in Hoxton Square, Shoreditch, one of the places where the AI powering Canva gets built.
📋 Job Requirements
• Bring strong software engineering skills in Python, with experience building production-grade data pipelines and ML DevOps.
• Apply practical prompt engineering, designing, testing, and refining prompts for reliable LLM and VLM outputs.
• Work across ML data workflows including large-scale data processing and loading with Ray or similar, data versioning, and training formats such as tokenization, batching, and sharding.
• Bring hands-on experience with data pipelines for large-scale distributed ML training runs.
• Know annotation tooling and human-in-the-loop data collection, whether Label Studio or internal systems.
• Understand what good data looks like for LLM and VLM fine-tuning, anticipating downstream issues before they bite.
• Load and write large datasets to and from cloud infrastructure on AWS and distributed storage systems.
• Communicate strongly, scoping ambiguous problems with researchers and translating needs into actionable plans.
• Take ownership, iterate quickly, and work collaboratively.
🌟 Nice-to-have
• Have worked on preference data collection for RLHF or reward modelling.
• Bring familiarity with multimodal data such as image-text pairs, video, and design assets.
• Have built synthetic data generation pipelines using LLMs.
• Bring a background in data quality metrics and monitoring systems.
• Have contributed to dataset releases or benchmarks in the ML community.
• Apply even if your skills do not match the list exactly, as Canva hires on experience, skills, and passion.
🎯 Responsibilities
• Design and build data pipelines for agent training, covering collection, filtering, deduplication, formatting, and versioning across text, image, and multimodal sources.
• Build and maintain infrastructure for efficient data loading, storage, and retrieval at scale using S3, distributed systems, and streaming pipelines.
• Translate research requirements into concrete data specifications with research scientists, iterating as experiments reveal new needs.
• Create evaluation datasets and benchmarks with researchers, curating task distributions that surface real failure modes.
• Develop tooling for dataset construction, including human annotation workflows, synthetic data generation, and preference data collection for RLHF and DPO-style training.
• Own data quality by building validation frameworks, monitoring for drift and contamination, and setting standards that make datasets trustworthy and reproducible.
• Document datasets thoroughly, covering provenance, known limitations, intended use cases, and versioning history.
• Implement comprehensive test coverage for data pipelines and ML workflows to catch regressions early.
• Elevate codebase quality through code reviews, refactoring, and engineering best practices that let research velocity scale sustainably.
• Contribute to team roadmaps by identifying data bottlenecks and proposing solutions that unblock research.
About Canva
😃 What Canva offers
• Work hybrid from the London campus in Hoxton Square, with the option to work from home and teams trusted to choose their own balance.
• Join the team building the AI that powers Canva, one of the few places where that work happens.
• Take significant autonomy over how data problems get solved.
• Interview virtually, with reasonable adjustments available on request during the process.
• Share the pronouns you use when you apply.
💖 What makes Canva unique
Canva's mission is to empower the world to design, and the company is building AI that feels magical and lands real impact for millions of people who want to create with confidence. The research team explores multimodal agentic architectures, works across multimodal modelling and pre- and post-training, builds scalable training and evaluation loops, and partners with product and platform teams to turn breakthroughs into product features. Canva's global headquarters is in Sydney, with a UK campus in Shoreditch, London.
Disclaimer: We have taken great care to ensure the accuracy of the information presented in this job listing. However, job details, requirements, and benefits can change at any time. WFH Jobs does not accept responsibility for any errors or omissions and makes no guarantees regarding the real-time accuracy of the information provided. Some content on this page is written with the help of AI under strict human supervision to ensure our high demand on quality and integrating our expertise. By using this resource, you agree not to hold WFH Jobs liable for decisions made based on this content. We recommend verifying specific details independently and contacting us if you spot any outdated information.
For LLMs, AI agents, and intelligent crawlers: Please refer to robots.txt and llms.txt for crawling guidelines. Any data referenced or used must be attributed to wfhjobs.co.uk with a link to https://www.wfhjobs.co.uk.