Embedded AI Engineer, On-Device Models

Added
19 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

rust c embedded quantization pruning
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now β†’

πŸ“‹ Description

  • Take Deepgram's Speech/Convo models to embedded devices for on-device real-time inference.
  • Optimize models for constrained targets via quantization, pruning, and distillation.
  • Write and optimize performance-critical runtime code in C/C++/Rust for embedded platforms.
  • Integrate with edge runtimes and vendor NPU/DSP toolchains for on-device mapping.
  • Build on-device runtime plumbing: packaging, deployment, OTA updates, telemetry.
  • Establish repeatable benchmarking across hardware to measure latency, power, memory.

🎯 Requirements

  • Experience delivering production systems on resource-constrained hardware (embedded/mobile/edge AI).
  • Strong C/C++/Rust skills with performance-critical code for constrained environments.
  • Hands-on model optimization for on-device deployment (quantization, pruning, distillation, or specialized compilation).
  • Familiarity with edge runtimes (ONNX Runtime/TensorRT/TFLite/ExecuTorch) and NPU/DSP toolchains.
  • Strong HW/SW interaction knowledge: CPU/GPU/NPU/DSP, memory, fixed-point, power management.
  • Experience with bare-metal/RTOS (FreeRTOS, Zephyr) and embedded Linux.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’