Member of Technical Staff, AI Training Infrastructure

Added
3 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

azure docker aws kubernetes gcp

๐Ÿ“‹ Description

  • Design scalable infrastructure for large-scale model training
  • Build and maintain distributed training pipelines for LLMs
  • Optimize training across GPUs, nodes, and data centers
  • Implement monitoring, logging, and debugging tools for training ops
  • Architect data storage for large-scale training datasets
  • Automate provisioning, scaling, and orchestration

๐ŸŽฏ Requirements

  • Bachelor's degree in CS/CE or equivalent practical experience
  • 3+ years in distributed systems and ML infrastructure
  • PyTorch experience
  • Cloud platforms: AWS, GCP, Azure
  • Kubernetes and Docker for containerization/orchestration
  • Master's or PhD in CS or related field (preferred)

๐ŸŽ Benefits

  • Comprehensive benefits package
  • Professional development and learning opportunities
  • Collaborative, fast-moving team
  • On-site work across New York and San Mateo
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’