Senior AI Infrastructure & Platform Operations Engineer

Added
22 minutes ago
Location
Type
Full time
Salary
Salary not provided

Related skills

linux grafana prometheus kubernetes opentelemetry
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now →

📋 Description

  • Lead investigations and resolve complex infrastructure incidents.
  • Serve as senior escalation point during critical service events.
  • Support large-scale NVIDIA GPU infrastructure and high-performance networking environments.
  • Troubleshoot Linux, Kubernetes, networking, storage, and hardware issues.
  • Analyze platform performance, capacity, stability, and reliability trends.
  • Lead root cause analysis and drive long-term fixes.

🎯 Requirements

  • 7+ years in infrastructure/Platform ops, SRE, or related roles.
  • Expert Linux administration and troubleshooting.
  • Strong networking; diagnose performance, connectivity, reliability issues.
  • Production Kubernetes experience.
  • Experience supporting large-scale production infrastructure.
  • Proven incident leadership and root-cause analysis.
  • Strong observability and monitoring practices.
  • Excellent troubleshooting, communication, and collaboration.

🎁 Benefits

  • Operate advanced AI infrastructure in production.
  • Work with NVIDIA GPUs, Kubernetes, and high‑perf networking.
  • Define reliability standards and operational practices.
  • Influence AI-powered capabilities via k0rdent AI.
  • Collaborate with engineers to resolve complex infra challenges.
  • Join a growing org investing in AI infrastructure.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →