Performance Engineer, Inference Engine

Added
7 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

optimization rust performance llm cuda
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now β†’

πŸ“‹ Description

  • Work on building and optimizing the inference engine at Anthropic scale
  • Improve throughput, cost, reliability, and latency across all accelerator and cloud platforms
  • Keep device utilization high so accelerators are not waiting due to overheads
  • Reuse model state cache instead of recomputing when cheaper
  • Measure, model, then change performance to add observability and improve efficiency
  • Ensure model quality and robustness

🎯 Requirements

  • Working mental model of LLM inference: how prefill and decode land on an accelerator's compute
  • Understanding of host responsibilities during inference
  • Proven quick learner in deep, unfamiliar systems
  • Strong systems programming skills in languages like C++, Rust
  • Analytical about performance: observe and profile first
  • Low ego and willingness to ask questions

🎁 Benefits

  • Competitive compensation and benefits
  • Optional equity donation matching
  • Generous vacation and parental leave
  • Flexible working hours

πŸ›ƒ Visa sponsorship

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’