Member of Technical Staff - ML Performance

Added
2 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

linux cuda tensorrt nvidia vllm

📋 Description

  • Build ML systems performant at scale
  • Contribute to open-source projects and Modal's runtime
  • Push language and diffusion models toward higher throughput
  • Improve latency and efficiency of GPU inference
  • Collaborate with research and engineering teams

🎯 Requirements

  • 5+ years of high-quality, high-performance code
  • Experience with torch, ML frameworks, and inference engines (vLLM or TensorRT)
  • Familiarity with Nvidia GPU architecture and CUDA
  • ML performance engineering: boost GPU perf, debug SM occupancy, reduce host overhead
  • Nice-to-have: Linux kernel, file systems, containers
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →