Staff Site Reliability Engineer

Added
5 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

datadog sre aws kubernetes go

πŸ“‹ Description

  • Define SLIs and SLOs for Fingerprint's critical request paths (identification, events, server APIs
  • Introduce error budgets as the mechanism for balancing reliability investment against feature work
  • Own the reliability metrics that leadership uses to judge progress; be the No Nonsense voice on
  • Strengthen the incident lifecycle end to end: detection, response, communication, postmortem
  • Close the "customers find out before we do" gap: drive alert quality, correctness anomaly
  • Lead reliability reviews for high-risk changes and new services (production readiness, capacity

🎯 Requirements

  • 10+ years of engineering experience, with 3+ years as an SRE, production engineer, or
  • Deep experience with SLI/SLO design and error budgets in practice, including the hard part: getting
  • Strong incident leadership: you have run incident response and postmortems for high-severity
  • Hands-on depth in distributed systems failure modes β€” cache/database saturation and cascading
  • Comfortable reading and writing production code (Go, TypeScript, or similar) and infrastructure as
  • Track record of leading through influence: you have changed how teams you did not manage operate
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’