Staff Engineer, Inference Optimizations

Added
8 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

ai hpc cuda triton rocm
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now →

📋 Description

  • Performance Architecture: Lead perf benchmarks for inference engine and GPU kernels.
  • Deep-Dive Optimization: Solve complex perf issues across memory, precision, multi-node GPUs.
  • Technological Innovation: Apply cutting-edge optimization tech to Gen AI.

🎯 Requirements

  • 5+ years in HPC/AI infra; track record solving bottlenecks.
  • Gen AI literacy: familiarity with Gen AI landscape and model families.
  • Optimization: attention-layer optimizations and distributed GPU parallelization.
  • Hardware fluency: NVIDIA/AMD GPUs and CUDA/ROCm.
  • Open source mastery: contributing to open-source projects.
  • Systems design: low-level GPU programming and memory access.

🎁 Benefits

  • Well-being: competitive benefits and employee support.
  • Career development: conferences, training, LinkedIn Learning.
  • Equity: salary plus bonus and equity grants.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →