Research Engineer - AI/RL Infrastructure

Added
6 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

kubernetes pytorch cuda ray flyte
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now →

📋 Description

  • Design and build training and evaluation infrastructure using large GPU clusters.
  • Build benchmarking and evaluation systems for model performance.
  • Develop data sampling, dataset generation, and data curation pipelines.
  • Enable high-throughput distributed training across cloud environments.
  • Collaborate with AI research/autonomy/platform teams to productionize research.

🎯 Requirements

  • Production ML systems across training, evaluation, data, deployment.
  • Experience building an ML platform for training/eval/deploy.
  • Performance engineering for large-scale ML training; profiling and optimization.
  • Systems-level debugging for distributed training.
  • Deep familiarity with open-source ML/systems; when to adopt vs build.
  • Technical experience: PyTorch, CUDA, Ray, Flyte, Kubernetes.

🎁 Benefits

  • Health, dental, vision, life and disability insurance
  • 401k with employer match
  • Learning and wellness stipends
  • Paid time off
  • Equal opportunity employer/notice provisions
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →