PhD Research Scientist Intern - Reinforcement Learning for Diffusion Modelling

Added
18 days ago
Type
Internship
Salary
Salary not provided

Related skills

pytorch rl lora ppo dpo

📋 Description

  • Designing and validating VLM-based evaluators for layered diffusion models.
  • Turning evaluators into reward functions for RL-based generative models.
  • Collaborating with research, engineering and product teams to move findings toward production.
  • Contributing to Canva's layered-generation roadmap and broader research community.

🎯 Requirements

  • Currently completing a PhD (ideal third year or later).
  • Strong diffusion or flow-matching background with policy-gradient RL for generative models (GRPO
  • Experience fine-tuning VLMs (e.g., LoRA) and designing prompts/rubrics for evaluation tasks.
  • Reward modelling, preference optimization, pseudo-labelling or distillation experience.
  • Ability to read a recent paper and reproduce it quickly; clear written and verbal communication.
  • Enjoy collaborating with researchers and engineers on hard problems and juggling multiple threads.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →