Staff Engineer, Distributed Storage and HPC & AI Infrastructure

Added
5 minutes ago
Type
Full time
Salary
Salary not provided

Related skills

python kubernetes go distributed storage ceph
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now →

📋 Description

  • Architect and implement the technical strategy and storage roadmap for Together AI, driving
  • Engineer and scale multi-petabyte AI/ML storage systems by integrating Vast, Weka, and Ceph while
  • Develop intelligent caching and tiered storage architectures to achieve extreme IOPS and
  • Tune storage isolation at the L2/L3 network layers to ensure secure, production-grade multi-tenancy
  • Code Kubernetes storage operators and controllers to enable automated provisioning, self-service
  • Engineer end-to-end data paths to achieve 10+ GB/s per GPU node; architect multi-tier caching for

🎯 Requirements

  • 8+ years in storage engineering, managing distributed storage at multi-petabyte scale
  • Proven track record deploying and operating high-performance storage for GPU/HPC clusters
  • Deep Kubernetes and cloud-native storage experience in production environments
  • Strong coding skills in Go and Python with demonstrated ability to build production-grade systems
  • BS/MS in Computer Science, Engineering, or equivalent practical experience
  • History of technical leadership: designing systems that significantly improved performance
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →