AI Infrastructure & Platform Operations Engineer (remote in the EU)

Added
less than a minute ago
Type
Full time
Salary
Salary not provided

Related skills

linux networking grafana prometheus kubernetes
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now โ†’

๐Ÿ“‹ Description

  • Monitor, operate, and support production AI infrastructure platforms.
  • Investigate and resolve infra, networking, hardware incidents.
  • Support NVIDIA GPU infrastructure and platform services.
  • Monitor and troubleshoot Kubernetes-based environments.
  • Collaborate with engineering teams, hardware vendors, and datacenters.
  • Participate in incident response and root cause analysis.

๐ŸŽฏ Requirements

  • 3+ years in infrastructure/platform/network ops, SRE, cloud, or datacenters.
  • Strong Linux administration and troubleshooting skills.
  • Good understanding of networking concepts and diagnosing infra issues.
  • Working knowledge of Kubernetes in production environments.
  • Experience supporting production infrastructure and services.
  • Strong analytical and problem-solving skills.
  • Experience with structured operational and incident management processes.
  • Ability to work in a shift-based operational environment.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’