Sr. Engineering Manager, AI Runtime

Added
19 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

pytorch deepspeed distributed training gpu nccl
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now β†’

πŸ“‹ Description

  • Lead and grow a high-performing engineering team for Custom Training infra.
  • Define and own the AIR product/technical roadmap for scalability.
  • Collaborate with product, research, platform, infra teams, and customers end-to-end.
  • Drive architectural decisions for managed GPU training at scale.
  • Advocate for customer needs to translate into product impact.
  • Build observability and reliability for long-running multi-node training.

🎯 Requirements

  • 8+ years of software engineering experience, with 3+ years in engineering management.
  • Track record building and operating GPU training infra at scale (100s/1000s GPUs).
  • Familiar with PyTorch, DeepSpeed, Megatron-LM and parallelism (FSDP, tensor/pipeline).
  • Experience with training resilience: checkpointing, elastic training, failure recovery.
  • Understanding GPU performance: NCCL, interconnects, memory optimization.
  • Experience building platform products with clear SLAs and customer ownership.

🎁 Benefits

  • Pay range transparency policy.
  • Comprehensive benefits and perks.
  • Strong commitment to diversity and inclusion.
  • Equal employment opportunity standards.
  • Region-specific benefits information and support.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’