Manager, Site Reliability Engineering

Added
5 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

datadog azure terraform github actions kubernetes

๐Ÿ“‹ Description

  • Lead and grow the SRE team; set on-call cadence and capacity planning.
  • Define SLOs/SLIs and error budgets; embed in release processes.
  • Lead major incident responses; conduct blameless post-incident reviews.
  • Champion proactive reliability via chaos engineering and game days.
  • Build AI-native operations: anomaly detection and automated triage.
  • Improve observability and platform reliability; manage on-call escalation.

๐ŸŽฏ Requirements

  • 10+ years in SRE/DevOps/Platform Eng.
  • 4+ years direct people management (remote/on-call teams).
  • Ownership of reliability outcomes; define SLOs/SLIs/error budgets.
  • Deep Azure experience (3+ yrs) with AKS, networking, identity; AWS/GCP ok.
  • Incident command experience; lead severity-1 incidents; blameless reviews.
  • Bachelor's degree in CS/IT/Engineering or equivalent.

๐ŸŽ Benefits

  • Free premium health, dental, life, and vision insurance.
  • Generous 401(k) match.
  • Paid sick leave; accrual policy per applicable law.
  • Company events, virtual happy hours and team-building activities.
  • Unlimited PTO and flexible vacation.
  • Virtual yoga, meditation or boot camp classes offered daily.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’