Staff Engineer, Inference Optimizations

Added
19 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

cuda tensorrt rocm openai triton fp8

📋 Description

  • Lead benchmarking and performance optimization for inference engine and GPU kernels.
  • Serve as IC leader to solve memory bandwidth and compute bottlenecks.
  • Guide the technical roadmap for high-performance inference workloads.
  • Architect decisions to maximize throughput and minimize latency for large models.
  • Mentor engineers via code and design reviews to raise the technical bar.
  • Collaborate with Product Management and TPMs to translate hardware limits into shippable features.

🎯 Requirements

  • 5+ years in HPC or AI infrastructure.
  • Gen AI literacy: LLM, VLM, LMM.
  • Optimization expert: attention layers and distributed GPU parallelism.
  • Hardware fluency: NVIDIA/AMD GPUs; CUDA, ROCm.
  • Open source mastery: experience building with and contributing to OSS.
  • Systems design: low-level GPU programming, memory access, parallel execution.

🎁 Benefits

  • Innovative, purpose-driven culture for builders.
  • Career development resources and training.
  • Comprehensive benefits and well-being programs.
  • Competitive salary, equity, and bonus structure.
  • Remote-friendly and flexible work options.
  • Inclusive, equal-opportunity employer.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →