Research Engineer / Performance Engineer, RL Distributed Systems

Added
1 hour ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

rust python kubernetes go rdma

πŸ“‹ Description

  • Design, build, and operate the distributed systems that run RL at scale, across training, sampling
  • Find and remove whatever currently limits the system, whether it's scheduling, data movement
  • Build fault tolerance into every layer: failure detection, isolation, and recovery that keep
  • Design resource management and autoscaling so that compute follows demand as a run's needs shift
  • Build observability that makes it possible to understand what a run is doing and why it slowed
  • Build automation that detects and remediates common problems, and design interfaces that let

🎯 Requirements

  • Strong software engineering skills in Python and at least one systems language such as Rust, C++
  • Experience designing, building, and operating large-scale distributed systems in production
  • Deep understanding of distributed systems fundamentals, including consistency, coordination
  • Ability to reason quantitatively about throughput, latency, and resource costs across compute
  • Experience debugging complex failures across many hosts and services, including failures you can't
  • Strong written communication, including design documents and incident writeups

🎁 Benefits

  • Competitive compensation and benefits
  • Optional equity donation matching
  • Generous vacation and parental leave
  • Flexible working hours
  • Lovely office space in which to collaborate with colleagues

πŸ›ƒ Visa sponsorship

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’