Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

Added
44 minutes ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

rust kubernetes go cuda tensorrt

πŸ“‹ Description

  • Own infrastructure to run training and inference workloads on GPU clusters.
  • Build a self-serve platform for inference engineers and researchers.
  • Operate the GPU fleet across providers with unified provisioning.
  • Develop scheduling, placement, and fault tolerance for GPUs.
  • Manage Kubernetes for GPU orchestration and multi-cluster ops.
  • Ensure reliability and observability of the platform.

🎯 Requirements

  • Deep Kubernetes experience with custom operators and CRDs.
  • Experience managing GPU clusters, NVIDIA hardware, CUDA, networking.
  • Experience across multiple clouds (CoreWeave, AWS, GCP or similar).
  • Strong distributed systems fundamentals: scheduling, resource allocation, fault tolerance.
  • Proficiency in Go, Rust or C++ for infra and systems code.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’