PhD Research Scientist Intern - Reinforcement Learning for Diffusion Modelling
CanvaViennaInternshipEngineering
Remotely
pytorchloramulti gpudporldiffusionppovlm
Job Description
📋 Description
- Designing and validating VLM-based evaluators for layered diffusion models.
- Turning evaluators into reward functions for RL-based generative models.
- Collaborating with research, engineering and product teams to move findings toward production.
- Contributing to Canva's layered-generation roadmap and broader research community.
🎯 Requirements
- Currently completing a PhD (ideal third year or later).
- Strong diffusion or flow-matching background with policy-gradient RL for generative models (GRPO
- Experience fine-tuning VLMs (e.g., LoRA) and designing prompts/rubrics for evaluation tasks.
- Reward modelling, preference optimization, pseudo-labelling or distillation experience.
- Ability to read a recent paper and reproduce it quickly; clear written and verbal communication.
- Enjoy collaborating with researchers and engineers on hard problems and juggling multiple threads.
Back to all jobs