Staff Site Reliability Engineer

Added
2 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

datadog aws kubernetes incident management site reliability engineering

πŸ“‹ Description

  • Define and implement SLIs and SLOs for critical production request paths
  • Introduce and champion error budgets as a framework for balancing reliability with feature delivery
  • Establish and maintain reliability metrics for engineering leadership
  • Strengthen the incident management lifecycle including detection, response, communication
  • Improve alert quality, anomaly detection, escalation processes, and shared operational tooling
  • Lead reliability assessments for high-risk changes and new services

🎯 Requirements

  • 10+ years of engineering experience, including at least 3 years in SRE, production engineering, or
  • Demonstrated experience owning reliability at a platform or organizational level
  • Deep practical experience designing and implementing SLIs, SLOs, and error budgets
  • Strong incident leadership experience, including high-severity incidents and postmortems
  • Advanced understanding of distributed-system failure modes
  • Strong hands-on experience with Kubernetes, AWS, and observability platforms like Datadog

🎁 Benefits

  • Fully remote working environment
  • Opportunity to become the first dedicated SRE and establish organization-wide reliability practices
  • High level of autonomy and direct influence over engineering standards and platform reliability
  • Opportunity to work across multiple engineering teams and critical production systems
  • Close collaboration with engineering leadership, architects, and technical leads
  • Opportunity to shape AI-assisted reliability practices
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’