Staff Engineer, Distributed Storage and HPC & AI Infrastructure

Added
22 minutes ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

terraform python kubernetes go ceph

πŸ“‹ Description

  • Architect and implement storage roadmap and high-performance architecture for scaling GPU fleet.
  • Scale multi-petabyte AI/ML storage; integrate Vast, Weka, Ceph; auto tiering.
  • Develop intelligent caching and tiered storage for extreme IOPS and GPU-scale throughput.
  • Tune storage isolation at L2/L3 networks for secure multi-tenant storage.
  • Code Kubernetes storage operators and controllers for automated provisioning and quotas.
  • Aim for 10+ GB/s per GPU node; design multi-tier caching; scale across thousands of nodes.

🎯 Requirements

  • 8+ years in storage engineering, managing distributed storage at multi-petabyte scale
  • Proven track record deploying and operating high-performance storage for GPU/HPC clusters
  • Deep Kubernetes and cloud-native storage experience in production environments
  • Strong coding skills in Go and Python with production-grade systems and tooling
  • BS/MS in Computer Science, Engineering, or equivalent practical experience
  • History of technical leadership: designing systems that significantly improved performance, reliability (99.999%+ uptime), or cost efficiency

🎁 Benefits

  • Competitive compensation, startup equity, health insurance, and benefits.
  • Flexible remote work options.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’