Staff Site Reliability Engineer

Added
42 minutes ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

datadog bash python kubernetes go
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now β†’

πŸ“‹ Description

  • Define the technical strategy for Observability, Alerting, and platform infra.
  • Lead reliable, scalable, secure cloud platforms and distributed systems.
  • Champion SLIs, SLOs, error budgets, capacity planning, readiness.
  • Lead complex production incidents and implement improvements.
  • Build self-service platform capabilities to reduce toil and boost velocity.
  • Mentor engineers and set reliability direction long-term.

🎯 Requirements

  • 12+ years in software engineering, infrastructure, platform engineering, or SRE.
  • 6+ years in SRE; 3+ years leading cross-functional initiatives for distributed production systems.
  • Expert in observability and platform infrastructure; incident response, capacity planning, automation.
  • Kubernetes experience; observability platforms such as New Relic or Datadog.
  • Proficient in Python, Go, Bash, or similar; production tooling/automation experience.
  • Mentor engineers and communicate technical risk clearly; regulated environments (SOC 2, PCI) preferred.

🎁 Benefits

  • Medical, Dental, & Vision Insurance
  • Competitive Pay and Fair compensation
  • Maternity & paternity leave for full-time employees
  • Short & long-term disability
  • Learn from a dedicated leadership team
  • Top-of-the-line company swag
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’