Added
16 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

python llm deepspeed distributed training profiling

šŸ“‹ Description

  • Optimize training and inference workloads to maximize throughput, minimize latency, and improve
  • Analyze and improve performance across the full technology stack, including GPU kernels, memory
  • Profile CPU, GPU, and distributed workloads to identify bottlenecks and use quantitative analysis
  • Design and implement performance improvements using Python, C++, and relevant AI systems
  • Optimize distributed training and inference architectures, including model parallelism
  • Evaluate and implement model compression techniques while carefully considering their impact on

šŸŽÆ Requirements

  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related technical
  • 6+ years of professional experience in performance engineering, machine learning systems
  • Strong programming proficiency in Python and C++, with the ability to develop production-quality
  • Hands-on experience optimizing deep learning workloads on modern GPU architectures.
  • Deep understanding of distributed training and inference techniques, including parallelism
  • Experience using profiling and instrumentation tools across CPU, GPU, and distributed environments.

šŸŽ Benefits

  • $100,000 annual salary for this full-time direct W2 position.
  • 100% remote work within the United States.
  • Opportunity to work on challenging AI optimization and high-performance computing problems.
  • Exposure to large-scale neural networks, modern GPU architectures, distributed systems, and
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →