AI Systems Research and Development Engineer – LLM Inference Systems & Optimization

Added
8 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

optimization ai llm cuda gpu

📋 Description

  • Design and develop high-performance LLM inference systems, spanning distributed serving, runtime
  • Develop novel techniques to improve inference latency, generation speed, throughput, memory
  • Explore advanced inference techniques including speculative and parallel decoding, prefill/decode
  • Develop adaptive and intelligent inference systems that automatically optimize execution for new
  • Apply AI-driven and AI-native approaches to systems engineering, including automated profiling
  • Independently identify high-impact performance and systems problems, formulate hypotheses

🎯 Requirements

  • Bachelor's degree in Computer Science, Computer Engineering, or related field; Master's degree or
  • 5+ years of experience in AI systems research, GPU kernel networking, experience with CUDA, Triton
  • Hands-on experience with modern LLM inference and serving frameworks such as vLLM, and TensorRT-LLM
  • Strong understanding of LLM architectures and distributed inference systems
  • Experience with performance-oriented libraries such as CUTLASS, cuBLAS, and related technologies
  • Ability to work independently, design novel solving algorithms, and communicate findings
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →