Senior AI Infrastructure  Platform Operations Engineer (remote in the EU)

Added
1 day ago
Type
Full time
Salary
Salary not provided

Related skills

linux grafana prometheus kubernetes elk

๐Ÿ“‹ Description

  • Lead technical operations for large-scale AI infrastructure with NVIDIA GPUs, Kubernetes, and
  • Escalation point for incidents and critical issues in production environments.
  • Shape reliability practices, automation initiatives, and platform evolution for AI-powered services.
  • Oversee platform performance, capacity, and reliability trends; drive long-term improvements.
  • Collaborate with engineering, hardware vendors, and datacenters to resolve challenges.
  • Participate in major incident management and service restoration.

๐ŸŽฏ Requirements

  • 7+ years in infrastructure/platform operations, SRE, or related roles.
  • Expert Linux administration and troubleshooting.
  • Strong networking and production Kubernetes experience.
  • Experience supporting large-scale distributed systems and incident leadership.
  • Proven root cause analysis and operational improvement skills.
  • Experience with observability/monitoring and reliability practices.

๐ŸŽ Benefits

  • Operate advanced AI infrastructure with latest NVIDIA GPUs and high-performance networking.
  • Influence operational standards and reliability practices for next-gen AI platforms.
  • Opportunity to contribute to k0rdent AI-powered capabilities.
  • Collaborate with skilled engineers on complex scale challenges.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’