AI Model Policy Trainer, Content Risk - Remote US

Added
15 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

ai trust & safety ml data annotation content moderation

πŸ“‹ Description

  • Evaluate user requests and AI model responses involving violence, weapons, threats, and dark
  • Distinguish fictional, educational, historical, and defensive violence from requests that seek
  • Assess whether a model's response gives meaningful real-world capability, regardless of how the
  • Distinguish expressions of anger, frustration, or dark humor from credible threats or crisis
  • Select the most defensible classification when a case is genuinely ambiguous, and write concise
  • Write and refine adversarial or borderline prompts that probe where a model draws the line

🎯 Requirements

  • You have spent serious time in violent fiction as a writer, game master, game designer
  • You have real-world exposure to violence and its consequences through military, law enforcement
  • You have worked with people in distress through crisis lines, counseling, threat assessment
  • You use AI tools heavily and have opinions about where they refuse too much, help too much, or miss
  • You notice when one word, contextual detail, or change in intent materially affects the answer
  • You can hold a strong opinion without becoming attached to being right
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Data Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Data Jobs

See more Data jobs β†’