Lead Machine Learning Engineer, Inference & Performance

Added
1 day ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

python kubernetes shell cuda gke
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now →

📋 Description

  • Optimize Inference: Build and tune production LLM serving with vLLM and SGLang—max throughput and
  • Profile & Accelerate Training: Instrument training runs; resolve bottlenecks with attention
  • Engineer for the Hardware: Understand GPU architecture and attention internals; pick the right
  • Serve at Scale: Deploy and operate multiple models within shared GPU clusters on GKE with
  • Drive Efficiency: Own GPU utilization metrics; improve throughput-per-dollar and fleet performance.
  • Collaborate & Consult: Work with clients to translate performance, latency, and cost

🎯 Requirements

  • Bachelor's or Master's in CS, Engineering, or related field
  • 5+ years ML/AI engineering with performance, infra, or systems focus
  • Proven production deployment/optimization of models
  • Experience profiling and improving GPU utilization for training or inference
  • Knowledge of Data Engineering and SQL

🎁 Benefits

  • Competitive salary and comprehensive benefits
  • Health Insurance, Paid Leave, Holidays, Sick Leave
  • Parental Leave, Bereavement Leave, 401(k) match, Employee referral bonuses
  • Benefits overview: https://egen.ai/people/#benefits
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →