Added
7 days ago
Type
Full time
Salary
Salary not provided

Related skills

golang docker python kubernetes pytorch
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now β†’

πŸ“‹ Description

  • Build and deploy production-grade LLM inference systems across GPU machines.
  • Design and operate model-serving infra with vLLM, SGLang, TensorRT-LLM.
  • Optimize latency, throughput, reliability, and cost.
  • Apply quantization, batching, caching, and routing for performance.
  • Develop infra in Python or Golang with scalable, maintainable code.
  • Own initial inference platform with CTO and evolve it as the org scales.

🎯 Requirements

  • Extensive experience building and operating production software or infra systems.
  • Experience deploying LLMs in production (vLLM, SGLang, TensorRT-LLM or similar).
  • Expertise optimizing workloads via quantization, batching, routing, etc.
  • Strong Python or Golang production-grade coding skills.
  • Understanding of production inference architectures from request to response.
  • Excellent problem-solving and independent diagnostic ability.

🎁 Benefits

  • Competitive package including equity.
  • Health, dental, vision, life insurance; dependents coverage where available.
  • Country-specific benefits.
  • Flexible, outcome-driven schedule and remote-first environment.
  • Global distributed team and significant ownership over architecture.
  • Direct collaboration with senior leadership and Product teams.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’