Machine Learning Engineer, Speech - Joint Audio-Video Modeling

Added
2 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

video pytorch transformers audio diffusion

πŸ“‹ Description

  • Own audio side of multimodal generation end-to-end.
  • Design/train audio VAEs, neural codecs, vocoders.
  • Build diffusion/flow-matching models for AV generation.
  • Design audio conditioning and cross-modal alignment in AV models.
  • Run experiments and evaluations; drive data quality and metrics.
  • Lead small projects and collaborate cross-team.

🎯 Requirements

  • Experience with large-scale audio models (>8B params).
  • Diffusion/flow-matching transformers; samplers, schedules, conditioning, distillation.
  • Training audio VAEs, neural codecs, vocoders; objectives.
  • Multi-node, multi-GPU distributed training (FSDP/DeepSpeed).
  • Strong software engineering; production-grade PyTorch code.
  • Shipped speech/audio or multimodal models to production.

🎁 Benefits

  • Competitive salary and equity.
  • Medical, dental, and vision insurance premiums covered.
  • 42 days paid time off annually.
  • 401(k) retirement plan.
  • Lifestyle spending account – $500/month.
  • In-office meals and One Medical membership, and more.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’