Added
7 days ago
Type
Full time
Salary
Salary not provided

Related skills

golang python kubernetes pytorch cuda
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now โ†’

๐Ÿ“‹ Description

  • Build and deploy production-grade LLM inference systems across GPU machines.
  • Design and operate model-serving infra using vLLM, SGLang, TensorRT-LLM.
  • Optimize workloads for latency, throughput, and cost.
  • Develop robust infra with Python or Golang for production use.

๐ŸŽฏ Requirements

  • Extensive experience building/operating production infra.
  • Production deployment of LLMs (vLLM, SGLang, TensorRT-LLM, etc.).
  • Expertise in quantization, batching, caching, routing techniques.
  • Strong Python or Golang production code skills.
  • Understanding of production inference architectures.
  • Problem-solving and independent investigation skills.

๐ŸŽ Benefits

  • Competitive compensation including equity.
  • Health, dental, vision, life insurance; dependents coverage where available.
  • Benefits adapted to country of employment.
  • Remote-friendly, flexible schedule and global team.
  • Ownership of architecture and roadmap for inference platform.
  • Collaboration with senior leadership and product teams.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’