Staff Engineer, Inference Optimizations

Added
1 day ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

cuda tensorrt triton rocm fp8
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now →

📋 Description

  • Performance architecture: Benchmark and optimize inference engine and GPU kernels.
  • Deep-dive optimization: Address bottlenecks in attention, memory, and parallelism.
  • Technological innovation: Apply cutting-edge optimization techniques for Gen AI workloads.
  • Hardware & ecosystem mastery: Advise on NVIDIA/AMD GPUs, CUDA, ROCm, and stacks.
  • Precision optimization: Develop FP8/INT8/BF16/FP4 quantization to boost throughput.
  • Technical mentorship: Lead code/design reviews and raise technical bar.

🎯 Requirements

  • Technical Depth: 5+ years in HPC/AI infra; solve compute and memory bottlenecks.
  • Gen AI Literacy: Deep familiarity with Gen AI landscape and model families.
  • Optimization Expert: Hands-on with attention-layer optimizations and distributed GPU parallelization.
  • Hardware Fluency: Deep understanding of NVIDIA/AMD GPUs and CUDA/ROCm stacks.
  • Open Source Mastery: Experience building with and contributing to open-source.
  • Systems Design: Low-level GPU programming, memory access patterns and parallel execution.

🎁 Benefits

  • Innovate with purpose; think big, bold, and act like an owner.
  • Prioritize career development with conferences, training, and LinkedIn Learning.
  • Care for well-being with competitive benefits and flexible time off.
  • Salary, bonus opportunities, equity, and Employee Stock Purchase Program.
  • DigitalOcean is an equal-opportunity employer; no discrimination.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →