Senior / Staff ML Ops Engineer

Added
3 minutes ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

docker terraform helm aws python

πŸ“‹ Description

  • Build and evolve our training infrastructure on Kubernetes with Infrastructure β€” GPU scheduling
  • Shape the developer-facing surface β€” CLIs, SDKs, job submission, templates, paved paths β€” designed
  • Shorten the inner loop. Time to first training run, edit-to-signal latency, local iteration before
  • Evangelize best-in-class tooling and frameworks. Track what the ecosystem is shipping, evaluate
  • Strengthen the data and artifact layer. Dataset versioning, sharding, and high-throughput loading
  • Turn one-off Python into durable tooling β€” tested, documented, observable libraries, CLIs, and

🎯 Requirements

  • 5+ years of software or infrastructure engineering, including tools or platforms used by other
  • Hands-on Kubernetes expertise β€” GPU scheduling, autoscaling, Helm or equivalent, networking
  • Excellent Python, and a track record of designing APIs and CLIs other people enjoy using.
  • Practical AWS depth: object storage at scale, IAM, GPU compute, networking, cost management, and
  • Distributed training in PyTorch (DDP, FSDP, or similar), plus experiment tracking and model
  • Fluency with containers, CI/CD, and modern build systems, including large monorepos.

🎁 Benefits

  • Competitive compensation and equity awards.
  • Health and Wellness benefits encompassing Medical, Dental and Vision coverage (for full-time
  • Unlimited Vacation.
  • Flexible hours and Work from Home support.
  • Daily drinks, snacks and catered meals (when in office).
  • Regularly scheduled team building activities and social events both on-site, off-site &amp
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’