Research Engineer – Benchmarking, Evals & Failure Analysis

Added
2 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

sql nosql machine learning apis cloud platforms

📋 Description

  • Benchmarking: Design benchmarks and metrics for tool use.
  • Evaluation systems: Build end-to-end LLM evaluation runs, dashboards, and reporting.
  • Failure analysis: Run systematic failure analysis; categorize failures and feed improvements.
  • Rubrics and evaluators: Create rubrics, automated evaluators, and scoring frameworks.
  • Data quality and usability: Quantify data usability and impact on benchmarks; guide data curation.
  • Cross-team collaboration and ownership: Work with researchers and data producers; own workflows.

🎯 Requirements

  • Applied research in model evaluation, benchmarking, and failure analysis.
  • Coding skills and hands-on ML model evaluation experience.
  • Strong data structures, algorithms, and backend systems knowledge.
  • Experience with APIs, SQL/NoSQL, and cloud platforms for eval results.
  • Ability to reason about model behavior and data quality from evals.
  • Willingness to work in person in San Francisco five days a week.

🎁 Benefits

  • Bi-annual bonus and equity grant (vested over 4 years).
  • Relocation bonus up to $15k; housing bonus up to $10k near office.
  • $1.5K monthly meals stipend; $200 monthly laundry reimbursement.
  • $200 monthly personal wellness reimbursement; Free Equinox membership.
  • Health, Dental, Vision insurance.
  • In-person work at SF/ NYC / London offices.

🚚 Relocation support

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →