AI Evaluation Engineer

Added
12 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

nlp speech qa data annotation benchmarking

📋 Description

  • Design validation for agentic NLP and speech workflows.
  • Build, run, and improve regression evaluations and A/B tests.
  • Co-own LLM-judge metric development, calibration, and prompt refinement.
  • Create, configure, and monitor data annotation jobs for evaluation data.
  • Develop QA tooling, notebooks, and pipeline components for scalability.
  • Investigate bugs, triage issues, and decide follow-ups with engineering.

🎯 Requirements

  • Bachelor's or Master's in CS, Software Eng, Computational Linguistics, or related field.
  • 3+ years in QA, test engineering, model evaluation, or ML quality for AI products.
  • Experience designing structured test strategies across manual and automated workflows.
  • Comfort with complex AI systems such as speech, NLP, LLM, or agentic products.
  • Experience with evaluation datasets, gold sets, or benchmark creation for AI.
  • Strong analytical skills for investigating failures and identifying quality patterns.

🎁 Benefits

  • Work at the center of AI transformation in business communications.
  • Build and ship agentic AI products redefining how companies operate.
  • AI amplifies every employee’s impact.
  • Competitive salary, comprehensive benefits, and growth opportunities.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →