Member of Technical Staff, Performance Optimization

Added
2 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

kubernetes pytorch cuda triton nsight

📋 Description

  • Own performance optimization from GPU kernels to multi-GPU systems.
  • Improve latency, throughput, memory, and compute efficiency.
  • Profile performance to detect GPU/kernel bottlenecks.
  • Implement low-level optimizations using CUDA, Triton, and tools.
  • Scale inference and training across multi-GPU, multi-node environments.
  • Collaborate with ML researchers to tune models for hardware efficiency.

🎯 Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
  • 5+ years of experience in performance optimization or high-performance computing systems.
  • Proficiency in CUDA or ROCm and GPU profiling tools (e.g., Nsight, nvprof, CUPTI).
  • Familiarity with PyTorch and performance-critical model execution.
  • Experience with distributed system debugging and optimization in multi-GPU environments.
  • Deep understanding of GPU architecture, parallel programming models, and compute kernels.

🎁 Benefits

  • Meaningful equity in a fast-growing startup.
  • Comprehensive benefits package.
  • Opportunity to work with bleeding-edge AI infrastructure.
  • Ownership and impact in a fast-paced, results-driven environment.
  • Learn from world-class engineers and researchers.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →