Senior Engineer, Inference Data Plane

Added
16 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

grpc golang python tensorrt ray serve

πŸ“‹ Description

  • Technical leadership: deliver data plane components for AI inference.
  • System design: architect scalable, multi-tenant AI inference systems.
  • Performance optimization: tensor parallelism and KV cache optimization.
  • Collaboration: partner with product, customers, and engineers to align roadmaps.
  • Distributed serving at scale: Kubernetes-native frameworks for MoE models.
  • Flow control & load balancing: manage inference load and autoscaling.

🎯 Requirements

  • AI/ML domain knowledge: host large language models with vLLM, SGLang, or TensorRT.
  • Inference frameworks: experience with llm-d, NVIDIA Dynamo, or Ray Serve.
  • Inference engine depth: hands-on with vLLM or alternatives and internals (batching, caching).
  • Distributed inference fluency: KV-cache locality and cross-pod KV transfer.
  • Architecture proficiency: knowledge of LLM architectures and optimization.
  • Software engineering: Go or Python, and gRPC.

🎁 Benefits

  • Career development resources and growth opportunities.
  • Well-being programs including EAP, meetups, and flexible time off.
  • Competitive compensation with bonus and equity, including ESPP.
  • DigitalOcean is an equal-opportunity employer.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’