Site Reliability Engineer

Added
8 days ago
Type
Full time
Salary
Salary not provided

Related skills

rust java bash c python

πŸ“‹ Description

  • Own monitoring architecture and signal quality: what we alert on, suppress, and trust
  • Provide SEV command support: technical incident leadership, bridge coordination with the NOC, and
  • Run blameless postmortems and drive corrective actions to closed, not filed
  • Lead cross-functional reliability projects spanning compute, network, storage, and facility signal
  • Build and maintain playbooks, run game days, and keep cross-discipline dependency maps current
  • Define error budgets and availability objectives at campus and service boundaries

🎯 Requirements

  • Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or related field
  • 5+ years of experience in site reliability, systems engineering, or large-scale production
  • Proven large-scale incident command experience and calm technical leadership on a bridge
  • Demonstrated monitoring and observability design at fleet or campus scale
  • Experience working across at least two of: compute, network, storage, power, and cooling /
  • Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’