Machine Learning Engineer - Inference

Added
25 minutes ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

python pytorch cuda triton vllm

πŸ“‹ Description

  • Design and build production systems powering the Together AI inference engine at scale.
  • Develop and optimize runtime inference services for large-scale AI apps.
  • Collaborate with researchers, engineers, PMs, and designers to bring new features.
  • Conduct design and code reviews to ensure high-quality standards.
  • Create services, tools, and docs to support the inference engine.
  • Implement robust, fault-tolerant data ingestion and processing.

🎯 Requirements

  • 3+ years of experience writing high-performance, production-quality code.
  • Proficiency with Python and PyTorch.
  • Experience building high-performance libraries and tooling.
  • Strong grasp of OS concepts: multi-threading, memory, networking, storage, performance.
  • Knowledge of CUDA/Triton programming.
  • Knowledge of AI inference systems like TGI, vLLM, TensorRT-LLM, Optimum.

🎁 Benefits

  • Startup equity, health insurance, and other benefits.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’