Incident Management & Resilience Lead

Added
5 days ago
Type
Full time
Salary
Salary not provided

Related skills

datadog terraform aws postgresql redis

📋 Description

  • Own incident management across engineering: definition, process, tooling, adherence.
  • In the moment, join high-impact incidents to drive clarity and resolution.
  • Between incidents, run postmortems and ensure followups become changes.
  • Work across product teams plus Security and Support to uphold a resilience bar.
  • Coach engineers on incident response, communication, and ownership.
  • Read code, inspect telemetry, and lead real investigations.

🎯 Requirements

  • Built or lifted incident management capability at scale.
  • Clear, earned POV on what good looks like.
  • Read code, inspect telemetry, and lead investigations—not just chair calls.
  • Calm under sustained incidents; keep the standard without bruising relationships.
  • Hold standards across teams you don’t manage.
  • Experience with observability tooling (Datadog, Honeycomb) and AWS.

🎁 Benefits

  • Ownership, not tickets: flat structure and high autonomy.
  • Remote, async work with real overlap across ANZ and US-Pacific.
  • Small team (~150 people): changes visible across the org.
  • Equal Opportunity Employer with accommodations available.
  • First-of-its-kind role shaping incident management at Buildkite.
  • Opportunity to scale incident response for production systems.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →