Added
8 days ago
Type
Full time
Salary
Salary not provided

Related skills

golang docker python kubernetes pytorch

πŸ“‹ Description

  • Build and deploy production-grade LLM inference systems across GPU machines
  • Design and operate model-serving infrastructure with vLLM, SGLang, TensorRT-LLM
  • Optimize workloads for latency, throughput, reliability, and cost
  • Apply quantization, batching, caching, and routing to boost inference performance
  • Develop robust production infra using Python or Golang
  • Collaborate with CTO and Product to evolve the inference platform and roadmap

🎯 Requirements

  • Significant experience building/operating production software or infra systems
  • Production deployment of large language models with vLLM, SGLang, TensorRT-LLM or similar
  • Expertise in quantization, batching, caching, routing
  • Strong Python or Golang production-ready programming skills
  • Understanding of production inference architectures and end-to-end user request flow
  • Excellent problem-solving and communication skills

🎁 Benefits

  • Equity and competitive compensation
  • Health/dental/vision/life insurance with dependents coverage where available
  • Flexible, outcome-oriented schedule
  • Remote-first with globally distributed team
  • Ownership over architecture and long-term roadmap
  • Collaboration with senior leadership and Product teams
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’