PhD Research Scientist Intern - Reinforcement Learning for Diffusion Modelling
CanvaJob Description
📋 Description Work with Canva’s AI team on a live, industry-scale project. Collaborate with researchers and engineers toward production. Research and develop per-layer VLM evaluation and rewards. Contribute to layered-generation roadmap and publication efforts. Experience hands-on with real data, infra, and deadlines. Flexible hybrid work with in-person and remote collaboration. 🎯 Requirements PhD in progress, ideally in 3rd year or later. Diffusion/flow-matching background with policy-gradient RL. Experience fine-tuning VLMs (LoRA) and evaluation prompts. Reward modelling, preference optimisation, pseudo-labelling. Ability to read and reproduce recent papers quickly. Strong written and verbal communication of technical work. 🎁 Benefits Internship starting September; 16 weeks, London hybrid. Work on products used by hundreds of millions of people. Mentors and research-to-production exposure. Publication opportunities and potential for patentable work.