Research Engineer (Reinforcement Learning)

Added
5 hours ago
Type
Full time
Salary
Salary not provided

Related skills

python gpus lora vllm trl

๐Ÿ“‹ Description

  • Build training environments, verifiers, and infra for post-training models.
  • Own the synthetic data pipeline from generation to QA and validation.
  • Run end-to-end training experiments and analyze results.
  • Design evaluations models must pass before prod releases.
  • Adapt open-weight foundation models for agent needs.
  • Develop behaviors for voice and text-based agents.

๐ŸŽฏ Requirements

  • Strong Python engineering and production systems.
  • ML model lifecycle: data to production experience.
  • Data-centric mindset: coverage, quality, leakage care.
  • Design robust rewards, verifiers, and evals.
  • Experience with GPUs and training constraints.
  • Judgment on when training is right vs simpler approach.

๐ŸŽ Benefits

  • Impact on fast-growing dev platform.
  • Small, experienced team with ownership.
  • Competitive salary and equity.
  • Health, dental, vision benefits.
  • Flexible vacation policy.
  • Remote-friendly environment.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’