Senior AI Infrastructure & Platform Operations Engineer (remote in the EU)

Added
25 days ago
Type
Full time
Salary
Salary not provided

Related skills

linux networking grafana prometheus kubernetes

๐Ÿ“‹ Description

  • Lead technical operations for large-scale AI infra with NVIDIA GPUs, Kubernetes, and
  • Serve as escalation point during critical incidents and drive reliability initiatives.
  • Shape operational standards, automation, and platform evolution with k0rdent AI.
  • Collaborate with data centers, vendors, and engineering teams to resolve complex challenges.
  • Participate in incident management and service restoration activities.

๐ŸŽฏ Requirements

  • 7+ years in infrastructure/platform operations, SRE, or related roles.
  • Expert Linux administration and troubleshooting.
  • Strong networking and production Kubernetes experience.
  • Experience supporting large-scale production infra and distributed systems.
  • Root cause analysis and long-term operational improvement experience.
  • Observability and monitoring expertise; strong communication skills.

๐ŸŽ Benefits

  • Opportunities to work with advanced AI infra and NVIDIA GPUs.
  • Collaborate with top engineers on scale challenges.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’