Model Policy Manager, Agentic Safety

Added
10 minutes ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

data analysis cybersecurity evaluation ai safety

πŸ“‹ Description

  • Identify vulnerabilities that emerge as models interact with tools, data, and external systems, and
  • Develop threat models and empirical frameworks for understanding harmful outcomes from misaligned
  • Build frameworks for understanding harmful outcomes arising from model misalignment.
  • Identify the underlying behaviors and system conditions that drive those outcomes.
  • Turn findings into policy frameworks, evaluation criteria, online measurement and safeguards.
  • Develop human data campaigns and gold sets to ground measurement and evaluation of emerging

🎯 Requirements

  • Brings a strong background in AI agent safety, privacy, security, cybersecurity, or adjacent
  • Has demonstrated interest in AI alignment and a strong understanding of the technical drivers of
  • Has enough technical fluency to work directly with evaluation and training data, understand what
  • Is comfortable working hands-on with model data and evaluation results, including inspecting
  • Uses empirical evidence to develop and refine safety policies and safeguards.
  • Can translate complex or ambiguous alignment risks into precise behavioral expectations and

🎁 Benefits

  • Offers Equity
  • Relocation support
  • Hybrid work model (3 days in office, optional work from home on Thursdays and Fridays)
  • Open-plan offices with height-adjustable desks, conference rooms, phone booths, well-stocked

🚚 Relocation support

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Operations Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Operations Jobs

See more Operations jobs β†’