Staff / Principal Machine Learning Engineer, Serving - USA

Added
2 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

kubernetes llm ray quantization vllm

πŸ“‹ Description

  • Build and optimize real-time ML serving systems.
  • Containerize models and deploy to production.
  • Profile and optimize latency and throughput on GPUs.
  • Collaborate with research to ship reliable APIs.
  • Lead end-to-end deployment from research to prod.
  • Maintain reliability and performance as core features.

🎯 Requirements

  • Inference optimization: vLLM, TRT-LLM
  • Model acceleration: quantization, distillation, caching, batching
  • High-performance systems: C++, CUDA, Rust, Python
  • Distributed systems: Kubernetes, Ray, multi-GPU/multi-node
  • Public work: open-source contributions to inference engines
  • Background: PhD in CS, Physics, Math, or equivalent

🎁 Benefits

  • Relocation assistance
  • Mountain View office
  • Bonus and equity
  • Benefits package
  • Flat, collaborative culture with fast iterations

🚚 Relocation support

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’