Staff Site Reliability Engineer

Added
1 day ago
Type
Full time
Salary
Salary not provided

Related skills

datadog cloud terraform python kubernetes

πŸ“‹ Description

  • Define and execute the tech strategy for observability and reliability.
  • Lead scalable, secure cloud-native platforms and distributed systems.
  • Establish SLIs, SLOs, error budgets, capacity planning, automation.
  • Mentor teams and improve production incident response processes.
  • Develop self-service platform capabilities to reduce toil.
  • Promote AI/ML for observability and automated remediation.

🎯 Requirements

  • 12+ years in software/infra/platform/SRE.
  • At least 6 years SRE, 3+ years leading cross-functional initiatives.
  • Deep observability, reliability, incident response, capacity management.
  • Container orchestration with Kubernetes; Datadog, New Relic, or similar.
  • Strong Python, Go, Bash or similar programming skills.
  • Experience building automation, tooling, IaC, platform services.

🎁 Benefits

  • Competitive compensation package.
  • Medical, dental, vision insurance.
  • Maternity/paternity leave programs.
  • Short/long-term disability coverage.
  • Opportunity to work in fast-growing tech; mentorship from leaders.
  • Professional development and continuous learning.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’