Added
4 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

linux prometheus kubernetes hpc cuda

πŸ“‹ Description

  • Serve as a senior escalation point, troubleshooting hardware/driver/kernel issues
  • Quickly distinguish hardware failures, driver issues, kernel problems, and misconfig
  • Proactively identify process/tooling/documentation gaps and fix them
  • Use AI tools to build scripts and small internal tools to close gaps
  • Perform root-cause analysis across distributed systems, clusters, and GPU infrastructure
  • Collaborate with engineering to turn recurring pain points into permanent fixes

🎯 Requirements

  • 3+ years of hands-on HPC experience in admin, support, or engineering
  • Very strong Linux system administration experience
  • Experience with HPC environments; Linux cluster admin; Kubernetes/Slurm
  • Strong coding and CI/CD experience; AI-assisted tooling
  • Proficiency with monitoring/logging tools: Prometheus, Grafana, Datadog
  • CUDA/NCCL/NVLink/GPUDirect RDMA; IB/RoCE networking

🎁 Benefits

  • Generous cash and equity compensation
  • Health, dental, and vision coverage for you and dependents
  • Wellness and commuter stipends for select roles
  • 401k plan with 2% company match (USA employees)
  • Flexible paid time off plan
  • Equal Opportunity Employer
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’