ML Engineer, Inference & Optimization

Added
1 day ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

cuda quantization nccl attention_optimization sequence_parallelism

📋 Description

  • Accelerate Inference: Lead and implement advanced inference acceleration techniques, including
  • Maximize GPU Parallelism: Engineer and optimize GPU strategies across tensor, sequence, and
  • Programming for Performance: Develop and optimize high-performance computing kernels and
  • Advance AI Deployment: Bring state-of-the-art videogen and large language models into production in
  • Improve Training Efficiency (Bonus): Contribute to improvements in model training speed, stability

🎯 Requirements

  • Experience: 5+ years engineering experience, with a strong track record in inference acceleration
  • Inference Mastery: Expertise in inference optimization, including quantization, attention
  • GPU & Parallelism: Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP
  • AI Domain Knowledge: Familiarity with video generation (videogen) models and large language models
  • Collaboration & Ownership: Strong cross-discipline communication, self-driven

🎁 Benefits

  • Competitive salary in the AI industry
  • Equity in a fast-growing startup shaping the future of AI
  • Comprehensive health benefits, monthly stipends, company retreats
  • A supportive and collaborative office culture—we’re all building and launching together
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →