ML Researcher - Image / Video Diffusion

Added
9 days ago
Type
Full time
Salary
Salary not provided

Related skills

pytorch fsdp sp cp fp8

πŸ“‹ Description

  • Train diffusion models for image and video generation on large GPU clusters.
  • Fully optimize and profile large distributed training runs across model architectures, kernels
  • Implement and improve various distributed training strategies including FSDP, CP, SP, TP, and EP.
  • Continuously improve model quality and reliability through data, model architecture, training
  • Debug distributed training errors and implement fault tolerance solutions, identifying bad GPU
  • Ablate different architecture, attention, optimizer, data, and algorithmic choices to reliably

🎯 Requirements

  • Proven track record in working with image or video models at scale (publications or open-source
  • Strong proficiency in PyTorch and understanding of its inner workings.
  • Strong background in distributed training paradigms such as FSDP, CP, SP, USP, TP, and EP. Knowing
  • Experience in profiling and debugging large distributed training. Being comfortable with analyzing
  • Good knowledge of low precision training / inference in FP8, NVFP4, and MXFP8.
  • Solid understanding of diffusion model training pipeline across pretraining, midtraining

🎁 Benefits

  • Team: Work alongside a world-class team building the future of AI creative tooling
  • Impact: Significant scope and company-wide impact
  • Competitive compensation: generous salary & equity packages
  • Health & wellness: 100% health & 99% dental/vision insurance premiums covered for
  • Time off: Flexible PTO policy
  • Financial planning: 401k with a 4% company-sponsored match

πŸ›ƒ Visa sponsorship

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’