Senior Software Engineer I - AI Inference Data Plane

Added
1 day ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

grpc python kubernetes go ray serve
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now β†’

πŸ“‹ Description

  • Lead end-to-end design and delivery of AI data plane components hosting large models.
  • Architect high-scale AI inference cloud for multi-tenant workloads.
  • Optimize distributed inference with tensor/data parallelism, KV cache, and routing.
  • Build Kubernetes-native distributed inference using frameworks like llm-d, Ray Serve, KServe.
  • Collaborate cross-functionally with PMs, customer teams, and other engineers.
  • Mentor junior engineers and uphold observability and SLOs for platform health.

🎯 Requirements

  • AI/ML domain knowledge hosting large models with vLLM, SGLang, or TensorRT.
  • Experience with distributed inference frameworks llm-d, NVIDIA Dynamo, or Ray Serve.
  • Hands-on with vLLM or similar engines, incl. batching, paged attention, prefix caching.
  • Understanding of cluster-scale serving: KV-cache locality, cross-pod KV transfer, routing.
  • Contributions to open-source projects such as vLLM, llm-d, or SGLang.
  • Software engineering in Go or Python, and familiarity with gRPC.

🎁 Benefits

  • Career development resources and growth opportunities.
  • LinkedIn Learning access to 10,000+ courses.
  • Competitive benefits including EAP, meetups, and flexible time off.
  • Open-source mindset and potential equity programs.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’