ML Researcher - Posttraining

Added
9 days ago
Type
Full time
Salary
Salary not provided

Related skills

pytorch fsdp ppo dpo grpo

πŸ“‹ Description

  • Finetune diffusion models at scale to improve image aesthetics and quality.
  • Implement posttraining techniques ranging from supervised finetuning, preference optimization
  • Design comprehensive eval suites and reward designs for the reinforcement learning stage focused on
  • Train custom VLM as reward models as part of our reward design.
  • Train custom LLMs for prompt expansion through finetuning and reinforcement learning.
  • Coordinate with data teams and partners to manage collection of preference data and model

🎯 Requirements

  • Proven work of posttraining diffusion models for image or video generation.
  • Experience with large-scale model training, inference, and optimization.
  • Strong understanding of both LLM and diffusion post training pipelines and algorithms such as PPO
  • Strong proficiency in PyTorch and understanding of its inner workings.
  • Strong background in distributed training paradigms such as FSDP, CP, SP, USP, TP, and EP. Knowing
  • Good knowledge of low precision training / inference in FP8, NVFP4, and MXFP8.

🎁 Benefits

  • Competitive compensation: generous salary & equity packages
  • Health & wellness: 100% health & 99% dental/vision insurance premiums covered for
  • Time off: Flexible PTO policy
  • Financial planning: 401k with a 4% company-sponsored match
  • Meals in the office: breakfast, lunch, dinner - you name it, we'll cover it
  • Transit: Ubers covered to & from the office

πŸ›ƒ Visa sponsorship

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’