Member of Technical Staff, Performance & Capacity

Added
3 days ago
Type
Full time
Salary
Salary not provided

Related skills

training gpu infiniband rdma h100

๐Ÿ“‹ Description

  • Measure the fleet with real workloads to guide capacity decisions
  • Own and defend the capacity model with clear assumptions
  • Manage workload mix and preemptible fractions over time
  • Ensure data path speed with parallel file systems and data locality for GPUs

๐ŸŽฏ Requirements

  • 5+ years with GPU and large-scale compute workloads
  • System-level GPU performance knowledge for heterogeneous GPUs
  • Deep understanding of AI training and inference performance
  • Experience owning capacity decisions tied to money
  • Familiarity with parallel file systems and high-speed networking (InfiniBand, RDMA)

๐ŸŽ Benefits

  • Competitive compensation and meaningful equity
  • Opportunity to work at AI-native infrastructure scale
  • Equal opportunity employer with diverse perspectives
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’