Added
1 day ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

data python evaluation llm prompts

๐Ÿ“‹ Description

  • Design and implement evaluation frameworks for AI systems deployed in customer environments.
  • Define metrics to reflect real user needs and business objectives.
  • Build and maintain Python-based golden tests and regression suites.
  • Develop offline and online evaluation pipelines integrated into system iteration loops.
  • Calibrate and operate LLM-based graders aligning automated judgments with human assessments.
  • Collaborate with engineers and domain experts to guide production system development.

๐ŸŽฏ Requirements

  • 2+ years of software engineering experience
  • Strong Python engineering skills; production-grade evaluation/experimentation pipelines
  • Experience with Evaluation-Driven or Experiment-Driven Development
  • Ability to translate human judgment into code (test cases, scoring, graders)
  • Systems-oriented mindset; understanding prompts, agents, data, and deployment
  • AI-native working style; use AI tools for testing and debugging

๐ŸŽ Benefits

  • Base salary $150Kโ€“$250K; equity and comprehensive benefits
  • 100% medical, dental, vision for employee and dependents
  • Flexible time off; retirement and financial planning benefits
  • Wellness benefits; in-office meals and snacks
  • Access to AI models and tools; ownership of high-impact projects
  • Hybrid collaboration with in-office presence in SF/NYC locations
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’