Added
1 day ago
Type
Full time
Salary
Salary not provided

Related skills

gitops sre terraform python kubernetes

πŸ“‹ Description

  • Establish a company-wide SLO/SLA system and define reliability metrics.
  • Build MTTD/MTTR measurement, set goals and optimize.
  • Develop automated fault detection, diagnosis and recovery loop.
  • Promote chaos engineering with regular fault drills.
  • Establish change risk control: canary releases and auto rollback.
  • Drive cost governance with data-driven visualization and optimization.

🎯 Requirements

  • 10+ years in infrastructure/SRE; 5+ years leading a team of 10+
  • Deep SRE: SLO/SLI, error budgets, toil, capacity, incident mgmt.
  • Experience with large-scale cloud FinOps: >$5M/year spend.
  • Proficient in IaC and automation: Terraform, Kubernetes, GitOps.
  • Programming in Go or Python to build ops/tools.
  • Multi-cloud: AWS + another cloud; compliance-ready.

🎁 Benefits

  • Study Growth Fund for professional development
  • Internal events and team-building activities
  • Global collaboration with an international team
  • Career advancement opportunities at a global company
  • Internal mobility for long-term growth
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’