AI Infrastructure & Platform Operations Engineer (remote in the US)

Added
10 hours ago
Type
Full time
Salary
Salary not provided

Related skills

sre linux networking kubernetes infrastructure as code

๐Ÿ“‹ Description

  • Monitor, operate, and support production AI infrastructure platforms.
  • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
  • Support NVIDIA GPU infrastructure and associated platform services.
  • Monitor and troubleshoot Kubernetes-based environments.
  • Investigate performance, availability, and reliability issues across infrastructure and platform components.
  • Collaborate with engineering teams, hardware vendors, and data center staff to resolve issues.

๐ŸŽฏ Requirements

  • 3+ years in infrastructure/platform/network ops, SRE, cloud, or related roles.
  • Strong Linux administration and troubleshooting skills.
  • Good understanding of networking concepts and diagnosing infra issues.
  • Working knowledge of Kubernetes in production environments.
  • Experience supporting production infra and services.
  • Strong analytical and problem-solving skills.

๐ŸŽ Benefits

  • Work with an established Silicon Valley leader in cloud infrastructure.
  • Work with passionate colleagues helping Fortune 500 and Global 2000 customers.
  • Be part of cutting-edge, open-source innovation.
  • Thrive in a high-energy environment with openness and growth.
  • Professional development and training.
  • Attend conferences and working groups.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’