Staff+ Site Reliability Engineer, Safeguards ML Infra

Added
15 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

aws gcp ci/cd incident response site reliability engineering

πŸ“‹ Description

  • Launch captain model releases: stand up, configure, and verify safeguards for every new model, and
  • Own the off-cycle deployment of new safety classifiers as they ship from research β€” canarying
  • Verify that the right safeguards are provably live on the right models across every deployment
  • Automate yourself out of last quarter's work: turn launch runbooks into tooling, hand-built checks
  • Build and maintain a safeguards registry with full provenance β€” what is running in production, on
  • Participate in on-call and operational-duty rotations covering service incidents, model

🎯 Requirements

  • 8+ years of industry software engineering or site reliability engineering experience.
  • Have owned production change management at scale β€” deploy pipelines, config management systems
  • Have run high-stakes releases: served as a launch captain, incident commander, or release owner for
  • Have meaningful on-call experience for production systems, including incident response and
  • Have a desire to close the gap where nobody has yet raised their hand, even if it requires manually
  • Have hands-on experience deploying and operating on cloud platforms (AWS, GCP) at scale.

🎁 Benefits

  • Competitive compensation and benefits
  • Optional equity donation matching
  • Generous vacation and parental leave
  • Flexible working hours
  • Lovely office space in which to collaborate with colleagues
  • Visa sponsorship available

πŸ›ƒ Visa sponsorship

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’