Senior Principal AI Engineer

Added
18 days ago
Type
Full time
Salary
Salary not provided

Related skills

deepspeed slurm infiniband rdma nccl
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now โ†’

๐Ÿ“‹ Description

  • Lead design and optimization of large-scale distributed AI training systems across GPU
  • Design, build, and operate distributed training for large neural networks (autoregressive
  • Optimize multi-node, multi-GPU execution for throughput and training efficiency.
  • Diagnose performance bottlenecks across compute, memory, storage, and networking.
  • Improve reliability with fault-tolerant architectures and recovery strategies.
  • Collaborate with research and ML teams to productionize training pipelines.

๐ŸŽฏ Requirements

  • Extensive experience building/operating distributed systems or ML infra.
  • Experience running large-scale GPU workloads in production.
  • Strong PyTorch Distributed experience.
  • Deep understanding of data/tensor/pipeline parallelism.
  • Knowledge of GPU networking (NCCL, RDMA, InfiniBand, NVLink).
  • Experience with Slurm, Kubernetes, Ray, RunAI.

๐ŸŽ Benefits

  • Competitive compensation based on experience.
  • Fully remote work opportunity within the United States.
  • Opportunity to work on cutting-edge AI infra and large-scale systems.
  • High-impact role shaping AI development.
  • Collaborative environment with engineers and leaders.
  • Chance to solve challenging distributed computing problems.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’