Senior Engineer, Inference Data Plane

Added
19 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

grpc python kubernetes go vllm

πŸ“‹ Description

  • Act as a technical leader, driving end-to-end design of data plane for AI models.
  • Architect high-scale, multi-tenant AI inference cloud designs with strong availability.
  • Implement and optimize distributed inference with tensor/data parallelism and KV cache.
  • Collaborate with Product, customers, and other teams to align roadmaps.
  • Build on Kubernetes frameworks llm-d, Dynamo, Ray Serve, KServe for MoE models.
  • Solve LLM-serving flow control and load balancing, latency, and fairness.

🎯 Requirements

  • AI/ML Domain: host large language or multimodal models with engines (vLLM, SGLang).
  • Inference Frameworks: distributed inference frameworks llm-d, Dynamo, Ray Serve.
  • Inference Engine Depth: experience with vLLM or alternatives; batching and prefix caching.
  • Distributed Inference Fluency: cluster-scale serving; KV-cache locality and cross-pod transfer.
  • Architecture Proficiency: LLM architectures and optimization (continuous batching, quantization).
  • Software Engineering: Go or Python proficiency and gRPC.

🎁 Benefits

  • Career development resources, including conference reimbursements and LinkedIn Learning.
  • Well-being programs, local meetups, and flexible time off.
  • Remote-friendly with global benefits and location-aware policies.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’