Research Engineer – Benchmarking

Added
21 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

sql nosql apis llms failure analysis

📋 Description

  • Benchmarking: Design, implement, and maintain benchmarks and metrics for tool use, agentic
  • Evaluation systems: Build and operate LLM evaluation systems end-to-end runs, scoring, dashboards
  • Failure analysis: Run systematic failure analysis on model outputs (e.g., wrong tool use, reasoning
  • Rubrics and evaluators: Create and refine rubrics, automated evaluators, and scoring frameworks
  • Data quality and usability: Quantify data usability, quality, and impact on key benchmarks; use
  • Cross-team collaboration: Work with AI researchers, applied AI teams, and data producers to align

🎯 Requirements

  • Strong applied research background, with focus on model evaluation, benchmarking, and/or failure
  • Strong coding skills and hands-on experience with ML models and evaluation code.
  • Solid grasp of data structures, algorithms, and backend systems.
  • Comfort with APIs, SQL/NoSQL, and cloud platforms for running and storing eval results.
  • Ability to reason about model behavior, experimental results, and data quality from evals and
  • Excitement to work in person in San Francisco five days a week in a high-intensity, high-ownership

🎁 Benefits

  • Bi-annual performance bonus structure
  • Generous equity grant vested over 4 years
  • Up to $15k Relocation bonus
  • $10K housing bonus (if you live within 0.5 miles of our office)
  • $1.5K monthly stipend for meals
  • Free Equinox membership

🚚 Relocation support

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →