Senior/Staff Site Reliability Engineer - Data Center

Added
14 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

sre kubernetes kvm datacenter clusterapi

πŸ“‹ Description

  • Designing, building, and operating PathAI's on-prem/cloud environment with a focus on reliability
  • Advance SRE best practices with an emphasis on users, monitoring, and automation.
  • Operate data center to support a rapidly growing ML team and integrate with existing cloud infra to
  • On-call rotations and incident response participation to maintain uptime and resilience.

🎯 Requirements

  • 8+ years of relevant experience.
  • Familiarity with modern datacenter networks and cross-layer operation.
  • Experience administering physical hardware stacks (iDRAC/IPMI/Nvidia UFM/Juniper Systems).
  • Some experience with virtualization/containerization platforms (e.g., EKS-Anywhere/ClusterAPI/KVM).
  • Opinionated about storage solutions for high-performance workloads (Quobyte/S3/FSx/EFS).
  • Automation mindset: scripting and configuration management (Ansible/RedFish).

🎁 Benefits

  • Remote work option available.
  • PathAI is an equal opportunity employer.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’