AI Infrastructure & Platform Operations Engineer (remote in the EU)

Added
10 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

linux networking grafana prometheus kubernetes

๐Ÿ“‹ Description

  • Monitor, operate, and support production AI infrastructure platforms.
  • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
  • Support NVIDIA GPU infrastructure and associated platform services.
  • Monitor and troubleshoot Kubernetes-based environments.
  • Investigate performance, availability, and reliability issues across infrastructure and platform components.
  • Collaborate with engineering teams, hardware vendors, datacenter personnel, and service delivery teams to resolve technical issues.

๐ŸŽฏ Requirements

  • 3+ years of experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or related technical roles.
  • Strong Linux administration and troubleshooting skills.
  • Good understanding of networking concepts and experience diagnosing infrastructure-related issues.
  • Working knowledge of Kubernetes in production environments.
  • Experience supporting production infrastructure and services.
  • Strong analytical and problem-solving skills.

๐ŸŽ Benefits

  • Work with advanced AI infrastructure in production.
  • Gain exposure to NVIDIA GPUs, Kubernetes, and high-performance networking.
  • Help define how next-gen AI infrastructure is operated and supported.
  • Shape the future of AI-powered operations with k0rdent AI.
  • Join a growing organization investing in AI infrastructure and platform services.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’