Member of Technical Staff (AI Inference Engineer)

Added
2 hours ago
Type
Full time
Salary
Salary not provided

Related skills

rust python cuda gpu triton
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now โ†’

๐Ÿ“‹ Description

  • We build and run the inference engine behind every Perplexity query and deploy dozens of model
  • Responsibilities include supporting transformer-based retrieval, text-generation, and multimodal
  • Migrate GPU kernels to CuTe DSL to run on GB200 today and be portable to Vera Rubin racks.
  • Develop a Rust-native internal inference server to handle growing traffic and reduce Python-related
  • Profile and optimize performance from network ingress to batching and kernel interleaving.
  • Build dashboards, alerts, and automated remediation to catch regressions and respond to production

๐ŸŽฏ Requirements

  • Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar).
  • Understanding of modern LLM architectures and production deployment.
  • Experience operating production distributed systems under real load.
  • Proficiency across languages and layers: Rust for serving runtime, Python for model code
  • Ability to read a research paper, implement a kernel, and debug a production incident within a week.
  • Self-directed, thrives in fast-moving environments.

๐ŸŽ Benefits

  • Equity may be part of total compensation.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’