Added
3 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

distributed systems transformers cuda triton ssms

πŸ“‹ Description

  • Design and build low-latency, scalable inference stack for foundation models.
  • Work with research and product engineers to ship fast, cost-effective products.
  • Design and build robust inference infrastructure and monitoring.
  • Have significant autonomy to shape products and AI across devices.
  • Support real-time multimodal intelligence in production.
  • Improve performance, reliability, and observability.

🎯 Requirements

  • Strong engineering skills, navigate complex codebases with maintainable code.
  • Experience building large-scale distributed systems with high performance.
  • Technical leadership to execute zero-to-one results amid ambiguity.
  • Background in inference pipelines with ML and generative models.
  • Experience implementing state-of-the-art ML models to applied problems.
  • Preferable: CUDA or Triton experience.

🎁 Benefits

  • Compensation: Competitive base salary with equity
  • Health Insurance: Fully covered medical, dental, vision
  • Parental Leave: 9 weeks paternity, 12 weeks maternity
  • 401(k) plan
  • Commuter Allowance
  • Flexible PTO

πŸ›ƒ Visa sponsorship

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’