Senior Staff / Principal Machine Learning Scientist, AI Inference & Optimization

Added
2 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

python lora quantization onnx runtime qlora

πŸ“‹ Description

  • Own the model inference path: quantization, KV-cache, batching.
  • Tune latency and memory on constrained hardware.
  • Build runtime for bounded AI tasks with production signals.
  • Collaborate with systems and backend engineers to ship end-to-end.
  • Push toward dynamic task generation and context compaction.

🎯 Requirements

  • 10+ years in tech; 4+ years ML/AI hands-on.
  • Fine-tune with LoRA/QLoRA; quantization (GGUF/AWQ/GPTQ).
  • Runtimes: vLLM, TensorRT-LLM, ONNX Runtime, llama.cpp.
  • Strong Python; C++ interop as needed.
  • Transformer internals: KV-cache, attention, batching, memory.
  • Agentic coding systems (Claude Code, Pi, Codex) experience.
  • Education: MS in CS/ML/EE; PhD preferred.

🎁 Benefits

  • High-impact ownership of a net-new product.
  • Cutting-edge inference tech: quantization, KV-cache, memory management.
  • Real production scale with live customer signals.
  • Comprehensive benefits, including health plan.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’