Member of Technical Staff - Research, Inference

Added
2 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

serverless llm autoscaling quantization fp8

๐Ÿ“‹ Description

  • Own end-to-end inference research bets across models
  • Develop speculative decoding and disaggregated prefill/decode
  • Work on quantization (FP8, INT4) and KV-cache management
  • Optimize memory management and autoscaling for serverless traffic
  • Collaborate with customers and Forward Deployed Engineers to deploy models

๐ŸŽฏ Requirements

  • Research-leaning or systems background in LLM inference
  • Fluency in the LLM serving stack from kernels to schedulers
  • Track record shipping research or systems others build on
  • Drive to take a research bet from idea to result
  • Ability to work in-person in NYC or San Francisco

๐ŸŽ Benefits

  • In-person collaboration in NYC or SF offices
  • Work with cutting-edge AI infrastructure
  • Collaborative, mission-driven team culture
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’