Member of Technical Staff, ML Engineer

Added
19 hours ago
Type
Full time
Salary
Salary not provided

Related skills

terraform pytorch training gpu ray

📋 Description

  • Own the training and inference infrastructure that Core AI depends on: distributed training jobs, GPU scheduling, and model-serving systems (vLLM, SGLang, or comparable) for both proprietary models and self-hosted inference.
  • Build the tools and abstractions AI researchers use to launch training runs, iterate on inference providers, and route workloads across models, so a researcher's time goes into the science instead of the plumbing.
  • Partner with Engineering on the shared platform: capacity planning, observability, and reliability for GPU and inference infrastructure, so training and serving hold up to the same production bar as everything else we ship.
  • Debug and harden the training and inference stack under real load. Egress failures, stalled retries, and routing edge cases are your problem to close, not someone else’s ticket.
  • Stay hands-on. You write the code, not just the design doc, and you are the first call when a training job stalls or an inference path breaks.

🎯 Requirements

  • Three or more years building and operating ML training or inference infrastructure in production, at a company that trains or serves models at meaningful scale.
  • Hands-on experience with distributed training (multi-GPU or multi-node, using PyTorch, Ray, or comparable) and model-serving systems (vLLM, SGLang, Triton, or comparable).
  • Strong software engineering fundamentals. You can build a service that other engineers and researchers depend on every day, not a script that worked once.
  • Enough ML fluency to work productively with AI researchers: you understand training loops, reward signals, and inference-time behavior well enough to debug them, even without designing the algorithms yourself.
  • Experience building internal platform tools such as training-as-a-service APIs, inference gateways, or job schedulers.
  • Background in GPU infrastructure, CUDA, or performance engineering for ML workloads.

🎁 Benefits

  • Competitive compensation and meaningful equity
  • Remote/hybrid consideration on a case-by-case basis
  • Opportunity to work with leading AI researchers on discovery
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →