Senior Site Reliability Engineer

Added
27 minutes ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

ansible terraform grafana prometheus python

πŸ“‹ Description

  • Build and operate monitoring and alerting for cluster health β€” fabric, GPU, power/thermal, and
  • Remotely deploy and configure large-scale HPC clusters for AI workloads using automation
  • Automate cluster lifecycle: OS, firmware, drivers, networking as code (Ansible, Terraform)
  • Create runbooks and automated remediations for common cluster failure modes
  • Troubleshoot and resolve cluster issues across InfiniBand/RoCE, NCCL, GPU-direct, fabric
  • Participate in on-call rotations and lead incident response for cluster-level problems

🎯 Requirements

  • 7+ years of experience in Site Reliability Engineering, HPC Engineering, DevOps, or similar
  • Strong understanding of modern AI infrastructure including GPU architectures
  • Solid Linux-based systems knowledge in distributed environments
  • Experience with InfiniBand (IB), RoCE, CLOS fabrics, 100GbE, NCCL
  • Proficiency in Python and Go; experience improving internal tooling
  • Experience with monitoring and alerting tools (Prometheus, Grafana, Clickhouse); automation with

🎁 Benefits

  • Generous cash and equity compensation
  • Health, dental, and vision coverage for you and dependents
  • Wellness and commuter stipends for select roles
  • 401k Plan with 2% company match (USA employees)
  • Flexible paid time off
  • Equal Opportunity Employer
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’