Senior Site Reliability Engineer (SRE)

Added
8 minutes ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

azure java terraform aws python

๐Ÿ“‹ Description

  • Lead efforts to improve system reliability, scalability, and performance across critical services
  • Define and implement SLIs/SLOs and error budgets, and use them to guide engineering priorities
  • Design and develop observability systems (metrics, logging, tracing, alerting) that produce
  • Lead complex incident response, acting as incident commander when needed
  • Conduct postmortems focused on systemic causes rather than individual fault, and ensure corrective
  • Identify and eliminate toil through automation, tooling, and improved workflows

๐ŸŽฏ Requirements

  • 6-10+ years of experience in SRE, infrastructure, or backend systems engineering
  • Demonstrated experience of owning reliability outcomes for complex, distributed systems
  • Strong experience with cloud infrastructure (AWS, GCP, or Azure) and production-scale systems
  • Deep understanding of observability, incident management, and system performance
  • Proficiency in at least one programming language (e.g., Go, Python, Java) with a focus on
  • Able to change how other teams work without having managerial authority over them

๐ŸŽ Benefits

  • Impactful Work
  • Dynamic Culture
  • Comprehensive Benefits (medical, dental, vision, 401(k), wellness benefits)
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’