Senior Principal Site Reliability Engineer

Added
22 days ago
Type
Full time
Salary
Salary not provided

Related skills

aws grafana prometheus kubernetes go
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now โ†’

๐Ÿ“‹ Description

  • Design and build an enterprise-grade chaos platform across multi-cluster (K8s + EC2) and
  • Develop core chaos capabilities: fault injection at pod/node/AZ level
  • Support latency, packet loss, partitions, and dependency timeout injection
  • Production safety: blast radius control, one-click Kill Switch, auto rollback
  • Integrate fault scenarios with monitoring and SLOs for automated feedback loop
  • Enable traffic and fault isolation for experiments

๐ŸŽฏ Requirements

  • 8+ years backend/infrastructure experience; 3+ yrs chaos/stability engineering
  • Hands-on with large-scale production fault injection under safety constraints
  • Expertise in Kubernetes fault injection (Chaos Mesh / Litmus / custom)
  • Go or another backend language; system architecture design
  • Deep understanding of distributed failure modes
  • Observability stack: Prometheus, Grafana, Thanos, OpenTelemetry

๐ŸŽ Benefits

  • Study Growth Fund for professional development
  • Internal events and team-building activities
  • Global collaboration with international colleagues
  • Career advancement opportunities
  • Internal mobility for long-term growth
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’