Added
8 days ago
Type
Full time
Salary
Salary not provided

Related skills

golang docker python kubernetes pytorch

πŸ“‹ Description

  • Build and deploy production-grade LLM inference systems across GPU machines
  • Design and operate model-serving infra using vLLM, SGLang, TensorRT-LLM
  • Optimize inference workloads for latency, throughput, reliability, cost
  • Apply quantization, batching, caching, routing to boost perf
  • Develop infra in Python or Golang with scalable engineering focus
  • Partner with CTO to shape initial inference platform and roadmap

🎯 Requirements

  • Significant experience building/operating production software or infra
  • Production deployment of large language models (vLLM, SGLang, TensorRT-LLM, or equivalent)
  • Optimization techniques: quantization, batching, caching, routing
  • Strong Python or Golang production code skills
  • Understand end-to-end inference architecture from request to response
  • Excellent problem-solving and independent investigation skills

🎁 Benefits

  • Competitive package with equity
  • Health/dental/vision/life insurances
  • Benefits tailored to country of employment
  • Remote-friendly with flexible schedule
  • Global distributed team with ownership and impact
  • Work on GPU/AI infra at scale and open-source tech
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’