Senior AI Infrastructure  Platform Operations Engineer (remote in the US)

Added
10 hours ago
Type
Full time
Salary
Salary not provided

Related skills

linux kubernetes monitoring opentelemetry gpu

๐Ÿ“‹ Description

  • Lead investigations and resolution of complex infra, networking, and platform incidents.
  • Senior escalation point for operational events during critical service impacts.
  • Support large-scale NVIDIA GPU infrastructure and high-performance networking.
  • Troubleshoot Linux, Kubernetes, storage, and hardware issues.
  • Analyze platform performance, capacity, stability, and reliability trends.
  • Lead root cause analysis and drive long-term corrective actions.

๐ŸŽฏ Requirements

  • 7+ years of experience in infrastructure operations, platform operations, SRE, or related roles.
  • Expert-level Linux administration and troubleshooting.
  • Strong networking expertise for diagnosing performance and reliability issues.
  • Production experience operating Kubernetes in production environments.
  • Experience leading technical investigations and incident management.
  • Strong observability, monitoring, and reliability practices.

๐ŸŽ Benefits

  • Operate advanced AI infrastructure environments in production today.
  • Work with NVIDIA GPUs, Kubernetes, and high-performance networking tech.
  • Help define operational standards and reliability practices for AI infra.
  • Influence the adoption of AI-powered operational capabilities via k0rdent AI.
  • Collaborate with highly skilled engineers solving complex infra challenges at scale.
  • Competitive compensation with strong benefits package.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’