Staff Cloud SRE – AI/ML Platform & GPU Compute

Added
19 days ago
Type
Full time
Salary
Salary not provided

Related skills

datadog azure aws grafana prometheus
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now →

📋 Description

  • Founding Cloud SRE role for Wayve AI cloud infra.
  • Own reliability of Model Dev Platform and GPU Compute envs.
  • Scale multi-tenant GPU fleets and training workflows.
  • Define production standards and automation frameworks.
  • Collaborate with ML and software teams for readiness.
  • Design for reliability, performance, and scale.

🎯 Requirements

  • Experience as SRE/Production Engineer for large cloud systems.
  • GPU-backed environments or large ML infra experience.
  • Production ML pipelines (MLOps) in production.
  • Strong Kubernetes, prod clusters experience.
  • Production workloads in AWS, GCP, or Azure.
  • Distributed systems expertise; compute-heavy workloads.

🎁 Benefits

  • Hybrid work policy: office and remote time.
  • London-based full-time role in our office.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →