Director, Site Reliability Engineering (Production Engineering)

Added
40 minutes ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

prometheus kubernetes elasticsearch kafka clickhouse

📋 Description

  • Take end-to-end operational accountability for the availability, latency, and performance of Tier-0
  • Enforce multi-nines SLAs, SLOs, and error budgets across services while driving engineering
  • Oversee capacity planning, resource forecasting, and efficiency efforts to ensure key
  • Implement high-signal alerting frameworks and actionable runbooks, minimizing alert fatigue while
  • Lead Production Engineering across India, driving operational rigor while coaching technical

🎯 Requirements

  • Foundational understanding of AI/ML technologies and experience leveraging, securing, or
  • 12+ years of engineering experience, including 5+ years managing engineering managers and
  • Proven track record of operating effectively within globally distributed engineering teams
  • Deep understanding of modern observability architectures—including distributed tracing
  • Hands-on background architecting, deploying, and supporting hyper-scale cloud-native platforms
  • Proven expertise in site reliability engineering principles, chaos testing, capacity planning, and

🎁 Benefits

  • Various health plans
  • Time off plans for vacation and sick time
  • Parental leave options
  • Retirement options
  • Education reimbursement
  • In-office perks, and more!
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →