Staff Site Reliability Engineer

Added
16 hours ago
Type
Full time
Salary
Salary not provided

Related skills

datadog sre bash python kubernetes

πŸ“‹ Description

  • Lead reliability strategy for production systems.
  • Define and execute strategy for observability, alerting, and platform infra.
  • Lead cloud platforms, distributed systems evolution.
  • Establish reliability frameworks: SLIs/SLOs, capacity, readiness.
  • Drive automation to reduce operational complexity.
  • Lead complex incident investigations and improvements.

🎯 Requirements

  • 12+ years in software/infra/SE/SRE.
  • 6+ years in SRE with 3+ years leading initiatives.
  • Expert in observability, incident response, automation.
  • Kubernetes experience.
  • Observability tools: New Relic, Datadog.
  • Strong programming: Python, Go, Bash.

🎁 Benefits

  • Competitive compensation.
  • Medical, dental, vision insurance.
  • Maternity/paternity leave.
  • Disability coverage.
  • Opportunity to influence reliability practices.
  • Environment focused on innovation and learning.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’