Staff Software Engineer, Inference

Added
5 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

python kubernetes go cuda torchserve

πŸ“‹ Description

  • Lead architecture, performance, and reliability for Kubernetes-native inference at massive scale.
  • Define cross-cutting design initiatives: routing, adaptive scheduling, GPU resource mgmt
  • Implement advanced inference optimizations: speculative decoding, KV-cache reuse, benchmarking
  • Aim for strict P99 SLAs with metrics-driven engineering and observability.
  • Collaborate across infrastructure boundaries and mentor senior/mid engineers.

🎯 Requirements

  • 8–12+ years in large-scale distributed systems or cloud platforms.
  • Proven track record leading cross-team technical initiatives at scale.
  • Strong coding in Go, Python, or C++.
  • Deep Kubernetes production-scale experience (orchestration, scheduling, service design).
  • Strong knowledge of networked systems, performance optimization, distributed design.
  • Hands-on inference systems experience: batching, caching, memory opt, mixed precision (BF16/FP8)

🎁 Benefits

  • Competitive salary with discretionary bonus and equity.
  • Comprehensive benefits package including health, dental, vision.
  • 401(k) with employer match and flexible PTO.
  • Casual work environment and growth opportunities at a fast-growing AI/cloud company.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’