Director, Site Operations

Added
9 minutes ago
Type
Full time
Salary
Salary not provided

Related skills

bash python jira site reliability engineering high-performance computing
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now β†’

πŸ“‹ Description

  • Own Cluster Uptime: Serve as extreme owner of node, rack, and cluster health across 5+ sites
  • Lead a Large Operations Organization: Direct a 250+ person team spanning site managers, shift
  • Drive Node and Rack Remediation: Ensure systematic recovery of failed nodes and racks through
  • Partner Across Functions: Coordinate with facilities operations to limit downtime from power and
  • Own Vendor Execution: Direct vendors through hardware rework and field operations so repairs
  • Lead Site Reliability Engineering: Own the SRE organization responsible for proactive cluster

🎯 Requirements

  • Bachelor's degree and 7+ years of experience working in a large scale operations with 5+ years
  • Proven ability to lead large, multi-site, 24/7 operations organizations in fast-paced
  • Deep expertise in server hardware, cluster reliability, and data center technologies, from
  • Experience supporting compute-heavy environments like AI, machine learning, or high-performance
  • A track record of owning uptime, Service Level Agreements, or reliability metrics for large compute
  • Experience leading site reliability engineering or equivalent reliability-focused teams, including
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Operations Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Operations Jobs

See more Operations jobs β†’