Engineering Manager, GPU Infrastructure

Added
14 days ago
Type
Full time
Salary
Salary not provided

Related skills

terraform grafana prometheus kubernetes pytorch

📋 Description

  • Lead and mentor a GPU infra engineering team, fostering excellence.
  • Define and execute GPU cluster roadmap for deployment and scaling.
  • Oversee topology-aware scheduling, fault detection, and perf optimization.
  • Collaborate with cloud providers to validate/deploy new GPU architectures.
  • Ensure reliability, scalability, and security across GPU infra.
  • Partner with AI researchers to translate needs into robust solutions.

🎯 Requirements

  • Experience leading engineering teams with technical mentorship.
  • Strong communication to explain complex technical concepts.
  • Ability to make data-driven decisions under pressure.
  • Experience working in remote, distributed teams.
  • Commitment to inclusive, collaborative team culture.
  • Deep ML/HPC infra expertise: GPU/TPU clusters, distributed training (PyTorch, TensorFlow, JAX).

🎁 Benefits

  • Weekly lunch stipend in local currency.
  • Full health and dental benefits, mental health budget.
  • Retirement matching (RRSP/401K) and pension.
  • 100% parental leave top-up up to 6 months.
  • Annual enrichment benefits and education stipend.
  • $500 home office stipend.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →