Staff ML Performance Engineer (Inference Optimisation)

Added
1 day ago
Type
Full time
Salary
Salary not provided

Related skills

python cuda tensorrt triton opencl

๐Ÿ“‹ Description

  • Profile and pinpoint bottlenecks across the full inference stack (model graph, compiler/runtime
  • Implement and validate optimisations in compilers, runtimes, and/or kernels (e.g. operator fusion
  • Build robust benchmarking and regression testing to ensure performance improvements hold across
  • Optimise for multiple targets (e.g. NVIDIA Orin/Thor, Qualcomm) and work with teams to support
  • Collaborate with model developers to influence architecture and training/deployment decisions that
  • Contribute to technical roadmaps and tooling and help raise the standard of performance engineering

๐ŸŽฏ Requirements

  • Proven experience improving performance in production systems with tight constraints (latency
  • Strong proficiency with at least one relevant stack/toolchain (e.g. TensorRT, CUDA, Qualcomm QNN
  • Comfort operating at multiple levels of abstraction โ€” from high-level model behaviour down to
  • Strong software engineering fundamentals (debugging, profiling, testing, and maintainable code).
  • Clear communicator and collaborative teammate; able to align multiple stakeholders on performance

๐ŸŽ Benefits

  • Exposure to embedded or edge deployment of ML models, including benchmarking on real devices and
  • Experience with NVIDIA and/or Qualcomm SoCs and performance tooling.
  • Python and C++ proficiency.
  • Experience mentoring others and/or driving technical direction in a small, fast-moving team.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’