LLM Inference Frameworks and Optimization Engineer

Added
27 minutes ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

python pytorch cuda tensorrt moe

πŸ“‹ Description

  • Design fault-tolerant, high-concurrency inference engine for text, image, and multimodal models.
  • Implement distributed inference strategies: MoE parallelism, tensor parallelism, and pipeline parallelism.
  • Apply CUDA graph, TensorRT graph optimizations, and torch.compile for speed.
  • Collaborate with hardware teams to identify bottlenecks and co-optimize GPU inference.
  • Work with AI researchers and infra engineers to optimize end-to-end model serving.

🎯 Requirements

  • 3+ years in deep learning inference, distributed systems, or HPC.
  • Familiar with at least one LLM inference framework (TensorRT-LLM, vLLM, SGLang, TGI).
  • Background in GPU programming (CUDA/Triton/TensorRT), compilers, quantization, and GPU cluster scheduling.
  • Deep understanding of KV cache systems like Mooncake, PagedAttention, or in-house variants.
  • Proficient in Python and C++/CUDA for high-performance DL inference.
  • Transformer/LLM/VLM/Diffusion model optimization; workload scheduling, CUDA graph, compiled kernels.

🎁 Benefits

  • Experience with RDMA/RoCE in large-scale data center networks.
  • Familiar with distributed file systems (3FS, HDFS, Ceph).
  • Familiar with Kubernetes and open-source distributed scheduling/orchestration.
  • Contributions to open-source deep learning inference projects.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’