Added
7 days ago
Type
Full time
Salary
Salary not provided

Related skills

golang docker python kubernetes pytorch
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now β†’

πŸ“‹ Description

  • Build and deploy production-grade LLM inference systems across GPUs
  • Design and operate model-serving infra with vLLM, SGLang, TensorRT-LLM
  • Optimize workloads for latency, throughput, and costs
  • Apply quantization, batching, caching, and routing techniques
  • Develop robust infra using Python or Golang with scalable code
  • Own initial inference platform with CTO collaboration

🎯 Requirements

  • Significant experience building/operating production infra
  • Production deployment of LLMs (vLLM, SGLang, TensorRT-LLM)
  • Expertise in quantization, batching, caching, and routing
  • Strong Python or Golang coding skills
  • Understanding of production inference architectures
  • Strong problem solving and independent investigation skills

🎁 Benefits

  • Competitive compensation including equity
  • Health benefits and dependents coverage where available
  • Country-specific benefits; remote-friendly schedule
  • Remote-first, globally distributed team
  • Ownership over architecture and roadmap
  • Direct collaboration with senior leadership and Product teams
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’