HPC Infrastructure Site Reliability Engineer

Added
3 hours ago
Type
Full time
Salary
Salary not provided

Related skills

linux ubuntu prometheus cuda nvidia

πŸ“‹ Description

  • Operate and improve high-density AI/HPC infrastructure in a 24/7 production environment
  • Participate in a 24x7x365 on-call rotation, incident response
  • Troubleshoot complex issues across compute, networking, storage, and orchestration layers in GPU-accelerated environments
  • Lead performance evaluation and acceptance of new HPC infrastructure before production
  • Drive continuous service improvement through automation, tooling, and process refinement
  • Build and maintain infrastructure automation and tooling (IaC and scripting)

🎯 Requirements

  • 8+ years in SRE/Infra in large-scale 24/7 production
  • 2–3+ years HPC/AI infra with GPU compute at scale
  • Strong Linux (Ubuntu) admin and production troubleshooting
  • Bare-metal infra and out-of-band tooling (IPMI/iLO/iDRAC/Redfish)
  • Solid networking incl. InfiniBand and RoCE
  • Automation/scripting (Bash, Python, Ansible) and IaC

🎁 Benefits

  • Work with cutting-edge GPU/AI infrastructure
  • Global, distributed team with learning culture
  • Flexible, globally connected workplace
  • Opportunity to grow with a scaling business
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’