Added
8 days ago
Type
Full time
Salary
Salary not provided

Related skills

golang docker python kubernetes pytorch

๐Ÿ“‹ Description

  • Build and deploy production-grade LLM inference systems across GPU machines
  • Design and operate model-serving infra using vLLM, SGLang, TensorRT-LLM
  • Optimize workloads for latency, throughput, reliability, and cost
  • Apply quantization, batching, caching, and routing techniques
  • Develop production code in Python or Golang with emphasis on scalability
  • Own initial inference platform and evolve it as the org scales

๐ŸŽฏ Requirements

  • Significant experience building/operating production software or infra systems
  • Production deployment of large language models, ideally with vLLM, SGLang, TensorRT-LLM
  • Experts in quantization, batching, caching, routing for inference
  • Strong Python or Golang coding skills for production-quality code
  • Understanding of production inference architectures from user request to served response
  • Excellent problem-solving and independent investigation skills

๐ŸŽ Benefits

  • Competitive compensation with equity
  • Health, dental, vision, life insurance, with dependents coverage where available
  • Benefits adapted to country of employment
  • Flexible, outcome-focused schedule; remote-friendly
  • Remote-first environment with distributed team
  • Ownership over architecture, implementation, and long-term roadmap
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’