Inference Performance Engineer

Added
21 minutes ago
Type
Full time
Salary
Salary not provided

Related skills

rust python go transformers cuda
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now β†’

πŸ“‹ Description

  • Build and improve the inference runtime
  • Design scheduling, batching, KV cache, and prefill
  • Implement low-precision kernels and speculative decoding
  • Drive throughput, latency, and cost per token
  • Collaborate with hardware teams on kernels and graph optimizations
  • Own the OpenAI-compatible API surface and serving protocol
  • Build benchmarking and profiling infrastructure

🎯 Requirements

  • BS in CS, EE, or related field or equivalent experience
  • Software: Rust, Go, Python, or C++
  • Understanding of concurrency, memory, and tail latency
  • Inference tech: transformers, attention, KV cache, batching, quantization
  • Experience with model serving frameworks: vLLM, TGI, SGLang, TensorRT-LLM, llama.cpp
  • GPU/ASIC programming: CUDA, ROCm, Triton
  • Experience with low-precision inference: FP8, FP4, INT4
  • Profiling and benchmarking with Nsight or perf

🎁 Benefits

  • Top-tier compensation structured to recognize and retain the best talent
  • Meaningful equity
  • Comprehensive medical, dental, vision, life, and disability insurance
  • Parental leave for all new parents, including adoptive and surrogate journeys
  • Flexible PTO
  • Paid Holidays
  • Relocation support

🚚 Relocation support

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’