Member of Technical Staff (AI Inference Engineer)

Added
30 minutes ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

rust python cuda gpu inference
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now β†’

πŸ“‹ Description

  • We build and run the inference engine behind every Perplexity query and deploy dozens of model
  • Our stack is Rust, Python, CUDA, and CuTe DSL.
  • Join a team focused on high-performance ML inference systems in production.

🎯 Requirements

  • Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar).
  • Understand modern LLM architectures and production deployment.
  • Experience building/operating production distributed systems under real load.
  • Comfort with Rust for serving runtime, Python for model code, CUDA/CuteDSL for kernels.
  • Ability to read research papers, implement kernels, and troubleshoot production incidents.
  • Self-directed in fast-moving environments.

🎁 Benefits

  • Equity
  • Competitive compensation with salary range disclosed
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’