Product Manager - AI Inference & Model Serving

Added
1 hour ago
Type
Full time
Salary
Salary not provided

Related skills

serverless triton dynamo vllm tensorrt-llm
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now β†’

πŸ“‹ Description

  • Own strategy, roadmap, and lifecycle for inference and model serving
  • Define serverless inference, dedicated endpoints, autoscaling, and routing
  • Manage KV cache and observability for production inference workloads
  • Collaborate with NeoClouds and enterprise teams on requirements and architecture
  • Align latency, throughput, utilization, and cost with measurable outcomes

🎯 Requirements

  • 7+ years in product management or senior technical AI/ML roles
  • Strong knowledge of production AI inference: model serving, autoscaling, observability
  • Able to reason about GPU, network, storage, orchestration trade-offs
  • Experience with runtimes: vLLM, SGLang, TensorRT-LLM, Dynamo, Triton
  • Comfortable in architecture reviews and technical-commercial discussions with platform teams

🎁 Benefits

  • Work with an established Silicon Valley leader in cloud infrastructure
  • Collaborate with passionate colleagues across Fortune 500 and Global 2000 customers
  • Be part of cutting-edge, open-source innovation
  • Thrive in a high-energy environment with openness, collaboration, risk-taking, and growth
  • Professional development and training
  • Attend conferences and working groups
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Product Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Product Jobs

See more Product jobs β†’