AI Infrastructure & Platform Operations Engineer

Added
4 minutes ago
Type
Full time
Salary
Salary not provided

Related skills

grafana prometheus kubernetes elk opentelemetry
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now β†’

πŸ“‹ Description

  • Monitor, operate, and support production AI infra platforms.
  • Investigate and resolve infra, networking, hardware, and platform incidents.
  • Support NVIDIA GPU infra and platform services.
  • Monitor and troubleshoot Kubernetes-based environments.
  • Drive improvements in monitoring, observability, automation, and processes.
  • Maintain operational docs, runbooks, and knowledge articles.

🎯 Requirements

  • 3+ years in infrastructure operations, platform ops, SRE, or related roles.
  • Strong Linux administration and troubleshooting skills.
  • Solid networking concepts; diagnose infra issues.
  • Production Kubernetes experience.
  • Experience supporting production infrastructure and incident management.
  • Excellent communication and collaboration; able to work in shift-based environments.

🎁 Benefits

  • Work with advanced AI infra environments in production.
  • Exposure to NVIDIA GPU tech, Kubernetes, and high-performance networking.
  • Help define how next-gen AI infrastructure is operated.
  • Join a growing organization investing in AI infra and platform services.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’