Senior Forward Deployed Engineer I (AI Inference)

Added
11 hours ago
Type
Full time
Salary
Salary not provided

Related skills

golang python kubernetes vllm tensorrt-llm

๐Ÿ“‹ Description

  • Embed with AI startups to deploy high-throughput LLM serving
  • Own end-to-end design of multi-tenant AI inference systems
  • Profile bottlenecks and optimize TTFT/TPOT in GPU clusters
  • Collaborate with customers and internal AI infrastructure teams
  • Architect distributed inference with Kubernetes-native tools
  • Deliver production-grade, low-latency serving solutions

๐ŸŽฏ Requirements

  • 6+ years in AI/ML systems or forward deployed engineering
  • Hands-on with vLLM, llm-d, SGLang, TensorRT-LLM
  • Proficient in Python or GoLang; know gRPC; Kubernetes
  • Experience with distributed inference, caching, and batching
  • Strong customer-facing communication and ownership
  • Willing to travel up to 30%

๐ŸŽ Benefits

  • Competitive compensation and equity
  • Career development resources and external training
  • Global benefits and flexible time off
  • Equal opportunity employer commitment
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’