Senior Engineer, Inference Data Plane

Added
1 day ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

grpc python kubernetes go tensorrt

πŸ“‹ Description

  • Technical Leadership: lead end-to-end data plane design for large AI models.
  • System Design: architect high-scale, multi-tenant AI inference cloud.
  • Performance Optimization: optimize distributed inference with tensor/data parallelism and KV cache.
  • Collaboration: work cross-functionally with PMs, customers, and other teams.
  • Distributed Serving: Kubernetes-native frameworks (llm-d, NVIDIA Dynamo, Ray Serve) for scalable inference.
  • Flow Control & Load Balancing: optimize queue depth, cache locality, autoscaling.

🎯 Requirements

  • Go or Python with gRPC for large-scale services.
  • AI/ML domain: host LLMs with engines like vLLM, SGLang, TensorRT.
  • Inference frameworks: llm-d, NVIDIA Dynamo, Ray Serve.
  • Inference engine depth: vLLM and alternatives incl. batching, prefix caching.
  • Upstream track record: merged contributions to vLLM, llm-d, SGLang.
  • Architecture proficiency: LLM architectures and optimization (batching, quantization).

🎁 Benefits

  • We innovate with purpose and empower ownership.
  • Career development: conferences, training, LinkedIn Learning.
  • Well-being: EAP, local meetups, flexible time off.
  • Competitive compensation with bonus and equity.
  • Inclusive, equal-opportunity employer.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’