Forward Deployed Engineer (Inference & Post-Training)

Added
9 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

quantization vllm speculative decoding tensorrt-llm sglang

πŸ“‹ Description

  • Inference Engine Optimization: Select, configure, and optimize inference engine based on hardware
  • Configuration & Performance Tuning: Develop configuration updates to win critical POCs
  • Post-Training & Fine-Tuning: Drive hands-on RL training runs and optimize system design; guide
  • Strategic Customer Alignment: Act as the primary technical point of contact for aligned strategic
  • Opinionated Onboarding: Establish direct alignment with strategic customers at onboarding; ensure
  • Product Feedback Loop: Directly influence our software and model roadmap by surfacing insights from

🎯 Requirements

  • Experience: 5+ years in a technical role, with a strong focus on inference systems, open-source LLM
  • Inference Engine Depth: Expert-level, hands-on experience with inference engines (e.g., vLLM
  • Inference Optimization: Deep knowledge of KV cache tuning, speculative decoding, tensor
  • Post-Training Knowledge: Hands-on experience with fine-tuning and post-training pipelines
  • Model Landscape Awareness: Broad knowledge of state-of-the-art open-source models and strong
  • Coding Proficiency: Strong Python skills; comfortable working in production environments

🎁 Benefits

  • Competitive compensation
  • Startup equity
  • Health insurance
  • Other benefits
  • Flexibility in terms of remote work
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’