Manager, Site Reliability Engineering

Added
9 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

sre security devops monitoring incident management

๐Ÿ“‹ Description

  • Lead Product Operations Assurance with a DevOps first principals approach to fleet monitoring
  • Build a highly automated operating model that enables the team to support a rapidly growing
  • Establish incident management processes, including severity definitions, escalation paths, incident
  • Coordinate cross-functional engineering response to complex issues spanning software, firmware
  • Develop diagnostic playbooks, troubleshooting procedures, and operational tooling that improve
  • Partner with Data, Analytics, and Cloud Applications teams to define the telemetry, dashboards

๐ŸŽฏ Requirements

  • 12+ years of experience supporting complex production, industrial, energy, infrastructure
  • Demonstrated experience leading production incident response, technical troubleshooting, escalation
  • Experience leading SRE, DevOps, or production operations teams in highly automated environments
  • Strong ability to diagnose and coordinate resolution of problems spanning software, firmware
  • Experience with operational monitoring, observability, alerting, remote diagnostics, reliability
  • Proven ability to lead cross-functional teams through high-priority technical issues, communicate

๐ŸŽ Benefits

  • Competitive salaries
  • Stock options
  • 100% coverage of medical, dental, and vision premiums for full-time employees
  • 80% of healthcare premiums for dependents covered
  • At least 12 weeks of paid leave for new parents (up to 20 weeks for birthing parents)
  • Generous vacation policies

๐Ÿšš Relocation support

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’