Senior Site Reliability Engineer, AI Platform

Added
23 minutes ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

azure aws kubernetes gcp observability

πŸ“‹ Description

  • Own and evolve production infrastructure supporting AI-related workloads and services at scale
  • Design and operate highly available Kubernetes-based platforms
  • Drive reliability through SLOs, observability, capacity planning and production guardrails
  • Lead complex production investigations and turn findings into durable architectural improvements
  • Improve shared infrastructure across networking, databases, service communication and compute
  • Build better CI/CD, progressive delivery, automation and developer experience

🎯 Requirements

  • Strong hands-on production experience with at least one major cloud provider: GCP, AWS or Azure
  • Strong experience designing and operating Kubernetes and cloud-native production systems at scale
  • Strong understanding of distributed systems, networking and reliability engineering
  • Experience operating business-critical systems with strong availability, scalability and
  • Ability to independently own ambiguous, cross-team technical problems and drive them to measurable
  • Strong automation mindset and ability to balance reliability, engineering velocity and cost

🎁 Benefits

  • Flexible workplace model with options for fully remote or hybrid-remote work
  • Global presence with offices in Paris, NYC, London, Sydney and Bucharest

πŸ›ƒ Visa sponsorship

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’