Site Reliability Engineer, Lead

Added
24 minutes ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

cloud infrastructure devops distributed systems observability incident management

πŸ“‹ Description

  • Drive incident response processes, retrospectives, and continuous improvement to prevent recurring
  • Harden software, develop playbooks, improve oncall metrics and post-mitigation workstreams, and
  • Lead the engineering analysis and response to novel, rare or unexpected events encountered during
  • Champion a culture of reliability engineering across the organization, including partnering with
  • Contribute to system software architecture to help improve its robustness and debuggability

🎯 Requirements

  • Master's degree or PhD in Computer Science, Engineering, or a related technical field
  • 10+ years of experience working on large-scale production software systems, especially machine
  • 10+ years on hands on coding experience with C++
  • Experience organizing and running oncall rotations
  • Excellent influence without authority skills
  • Organizational awareness, extremely collaborative, strong communication, focus on value add to the

🎁 Benefits

  • Discretionary annual bonus program
  • Equity incentive plan
  • Generous Company benefits program
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’