Staff Engineer, Inference Optimizations

Added
1 day ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

cuda tensorrt triton rocm openai triton
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now →

📋 Description

  • Lead benchmarking and performance optimization for inference.
  • Engineer memory, precision management, and multi-node GPU parallelism.
  • Apply cutting-edge optimization techniques to Gen AI workloads.
  • Advise on hardware/software stacks (CUDA, ROCm, TensorRT) and procurement.
  • Mentor engineers through code reviews and design guidance.
  • Partner with Product and TPMs to translate hardware limits into features.

🎯 Requirements

  • 5+ years in HPC or AI infrastructure.
  • Gen AI literacy across LLMs, VLMs, and LMMs.
  • Optimization: attention layers and distributed GPU parallelization.
  • Hardware fluency: NVIDIA/AMD GPUs, CUDA/ROCm.
  • Open-source experience building or contributing to OSS projects.
  • Systems design: low-level GPU programming and memory access patterns.

🎁 Benefits

  • Career development and access to learning resources.
  • Competitive benefits and flexible time off.
  • Support for conferences, training, and education.
  • Inclusive, equal-opportunity workplace.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →