Site Reliability Engineer, AI Platform

Added
23 minutes ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

azure aws networking kubernetes gcp

๐Ÿ“‹ Description

  • Build and operate production infrastructure supporting AI-related workloads and services
  • Operate and improve highly available Kubernetes-based platforms
  • Improve reliability through SLOs, observability, alerting and capacity management
  • Investigate production issues and turn findings into durable fixes and improvements
  • Work across networking, databases, compute and service infrastructure
  • Improve CI/CD pipelines, deployment automation and developer experience

๐ŸŽฏ Requirements

  • Solid hands-on Kubernetes knowledge, including workloads, resource management, and production
  • Strong experience with Infrastructure as Code, and the lifecycle of cloud infrastructure
  • Solid experience building and operating CI/CD pipelines and automated deployment workflows
  • Hands-on experience with at least one major cloud provider: GCP, AWS or Azure
  • Good understanding of networking, distributed systems and reliability engineering
  • Experience with monitoring, observability and troubleshooting production systems
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’