Inference Infrastructure Architect

Added
4 minutes ago
Type
Full time
Salary
Salary not provided

Related skills

grafana prometheus python kubernetes opentelemetry
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now β†’

πŸ“‹ Description

  • Operate and expand the fleet efficiently, maximizing useful inference throughput per GPU-dollar
  • Make the product run on it, including serverless inference for the open-weight catalog, and
  • Build serverless serving pools, the fleet layer, the platform under it, weight logistics and

🎯 Requirements

  • Owned production LLM serving under meaningful traffic and latency constraints; scale on the order
  • Kubernetes on GPU fleets, operated end to end: GPU Operator, device plugins, node pools and
  • Deep operational command of vLLM or SGLang: deploying, tuning and upgrading it (parallelism with TP
  • Performance engineering at the system level: you read engine and DCGM metrics, reason from
  • Python and Go for automation; real Linux, networking and storage depth.
  • Work with the community in English and Chinese, and write runbooks and design docs people actually

🎁 Benefits

  • The newest hardware, ours from the metal up, with a globally expanding B300 fleet that Telnyx
  • Own the platform, build the team, with a greenfield inference platform where you set the pattern
  • Open-source first, with upstream contribution as part of the job and conference travel supported.
  • Based in mainland China, no relocation required, working remotely with a global, async-friendly

πŸ›ƒ Visa sponsorship

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’