Staff / Principal Machine Learning Engineer, Serving - Switzerland

Added
2 hours ago
Type
Full time
Salary
Salary not provided

Related skills

kubernetes llm ray quantization vllm

πŸ“‹ Description

  • Lead real-time inference optimization for serving
  • Implement model acceleration: quantization, distillation
  • Optimize performance on GPUs (C++, CUDA, Rust, Python)
  • Build scalable distributed systems (Kubernetes, Ray)
  • Own end-to-end deployment from research to prod
  • Collaborate with cross-geo teams and leadership

🎯 Requirements

  • Inference optimization: vLLM or TRT-LLM
  • Model acceleration: quantization, distillation, caching, batching, paged attention
  • High-performance systems: C++, CUDA, Rust, Python; GPU profiling
  • Distributed systems: Kubernetes, Ray, load balancing, multi-GPU/multi-node
  • Public work: OSS contributions to major inference engines
  • Background: PhD in CS/Physics/Math or equivalent experience

🎁 Benefits

  • Flat structure, fast iterations
  • Open-source contributions encouraged
  • Impact-focused culture with visibility of work
  • Remote work within Switzerland
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’