Staff Site Reliability Engineer (AI Platform)

Added
7 days ago
Type
Full time
Salary
Salary not provided

Related skills

terraform aws prometheus kubernetes opentelemetry

๐Ÿ“‹ Description

  • Own reliability and performance of AI infrastructure: gateway and inference.
  • Design/evolve AI Gateway: routing, failover, rate limiting, caching.
  • Build observability for AI systems: latency, throughput, SLOs, metrics.
  • Drive cost optimization and FinOps for AI workloads.
  • Run capacity planning and incident response; write runbooks.
  • Scale AI expertise across the org; set standards and coach teams.

๐ŸŽฏ Requirements

  • 5+ years in SRE/platform/infrastructure with production ownership.
  • Hands-on with LLM-backed systems in production (Bedrock, OpenAI, Anthropic).
  • Cloud-native: AWS, Kubernetes, Terraform/IaC, CI/CD.
  • Observability: Prometheus/Grafana, OpenTelemetry, define SLOs.
  • Proven cost-optimization: cut cloud/inference spend and visibility.
  • Staff-level influence: set technical direction beyond your team.

๐ŸŽ Benefits

  • Hybrid onboarding to start remote work; relocation support.
  • Comprehensive health insurance for you and family.
  • Professional development budget for conferences and courses.
  • Flexible benefits budget; you pick perks for health, home office, etc.
  • Hybrid work and flexible time off, planned with your team.
  • In-office perks: free meals and snacks.

๐Ÿšš Relocation support

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’