Staff / Principal Machine Learning Engineer, Serving - UK

Added
2 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

rust kubernetes llm cuda ray

📋 Description

  • Build real-time agentic ML serving systems.
  • Optimize inference with vLLM and TRT-LLM.
  • Develop high-performance GPU code: C++, CUDA, Rust.
  • Scale multi-GPU/multi-node inference (Kubernetes, Ray).
  • Own models from research to production.
  • Contribute to open-source projects and docs.

🎯 Requirements

  • Inference optimization with vLLM and TRT-LLM.
  • Model acceleration: quantization, distillation, caching, batching.
  • High-performance code: C++, CUDA, Rust, or optimized Python; GPU profiling.
  • Distributed systems: Kubernetes, Ray, multi-GPU/multi-node inference.
  • Full-cycle ownership: containerize models to production.
  • PhD or equivalent experience in backend/ML systems.

🎁 Benefits

  • Equity and benefits included.
  • Flat structure with fast iterations.
  • Visible work and open-source contributions.
  • Potential relocation support for US move.
  • Work with leading researchers in realtime voice ML.
  • Impactful, production-grade ML systems.

🚚 Relocation support

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →