Principal ML Ops Engineer (EMEA Remote)

Added
27 days ago
Type
Full time
Salary
Salary not provided

Related skills

terraform helm python kubeflow triton

๐Ÿ“‹ Description

  • Build and operate prod-grade model serving infra (vLLM, TGI, Triton)
  • Design deployment pipelines with blue/green and canary rollouts
  • Develop auto-scaling systems and multi-model serving architectures
  • Optimize GPU utilization, memory, and network throughput
  • Design observability for latency, throughput, GPU usage, and cost
  • Manage model registries and CI/CD for automated deployments

๐ŸŽฏ Requirements

  • 4+ years in ML Ops, Platform Eng, SRE, or similar ML infra roles
  • Hands-on with model serving frameworks such as vLLM, TGI, Triton
  • Strong container orchestration background and GPU-based production workloads
  • Experience with MLOps tooling (registries, experiment tracking, pipelines)
  • Python proficiency and IaC skills (Terraform, Helm, or similar)
  • Strong understanding of distributed systems and reliability engineering

๐ŸŽ Benefits

  • Take ownership of infrastructure powering a rapidly scaling AI cloud platform
  • Build ML inference systems from the ground up in a high-growth startup
  • Work at the intersection of distributed systems and GPU computing
  • Gain deep expertise in next-gen AI infrastructure and model serving
  • Influence core engineering decisions and scalable best practices
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’