Member of Technical Staff (AI Infrastructure Engineer)

Added
1 day ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

aws kubernetes pytorch distributed training hpc

πŸ“‹ Description

  • Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training
  • Manage and optimize Slurm-based HPC environments for distributed training of large language models
  • Develop robust APIs and orchestration systems for both training pipelines and inference services
  • Implement resource scheduling and job management systems across heterogeneous compute environments
  • Benchmark system performance, diagnose bottlenecks, and implement improvements across both training
  • Build monitoring, alerting, and observability solutions tailored to ML workloads running on

🎯 Requirements

  • Expert-level Kubernetes administration and YAML configuration management
  • Proficiency with Slurm job scheduling, resource management, and cluster configuration
  • Experience deploying and managing distributed training systems at scale (PyTorch)
  • Deep understanding of container orchestration and distributed systems architecture
  • Familiarity with LLM concepts and training processes; GPU resource optimization
  • Experience operating large-scale Kubernetes deployments in production
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’