Research Engineer - Inference

Added
8 hours ago
Type
Full time
Salary
Salary not provided

Related skills

cuda tensorrt gpu inference triton

๐Ÿ“‹ Description

  • Deploy state-of-the-art models to production
  • Own path from research checkpoint to serving infra
  • Optimize inference latency, throughput, and cost
  • Use quantization, distillation, KV-cache, batching, custom kernels
  • Build high-performance serving systems for real-time workloads
  • Create tooling to ship models quickly and safely

๐ŸŽฏ Requirements

  • Experience deploying and serving ML models in production
  • GPU programming and inference optimization: CUDA, Triton, TensorRT
  • Experience with vLLM, SGLang or similar serving frameworks
  • Ability to profile, diagnose, and eliminate bottlenecks

๐ŸŽ Benefits

  • Innovative culture with impact-focused work
  • Growth paths and opportunities to drive impact
  • Learning & development stipend
  • Social travel stipend for annual meetups
  • Annual company offsite in new locations
  • Monthly co-working stipend
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’