Staff Platform & Reliability Engineer

Added
7 minutes ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

gitops terraform aws kubernetes iso 27001

πŸ“‹ Description

  • Define customer-facing SLIs and SLOs across product surfaces, run an error-budget program that
  • Own the DR strategy: regional failover, written RTO/RPO per tier, resilience against third-party
  • Own a GitOps deploy path with progressive delivery, automated analysis, one-click rollback for
  • Own the AWS foundation, fully managed as code; the Kubernetes platform and service mesh; capacity
  • Manage incidents end-to-end: paging and severity policy, incident-commander rotation, status-page
  • Build observability for metrics, logging, and distributed tracing; dashboards and alerts as code

🎯 Requirements

  • Senior-most, hands-on IC with production uptime experience and on-call ownership.
  • Deep production Kubernetes on AWS including service mesh, with GitOps-based delivery across many
  • Infrastructure-as-code at multi-account scale, including reshaping a large existing estate.
  • Built an SLO/error-budget practice that influenced release decisions with burn-rate alerting.
  • Delivered multi-region or DR capability with defined RTO/RPO and proven failover tests.
  • Security as daily practice: least-privilege IAM, secrets management, admission and network policy

🎁 Benefits

  • 100% paid health, dental & vision care
  • 401(k) & financial wellness perks
  • Daily meals on us
  • Commuter benefit
  • Monthly wellness stipend
  • Claude Enterprise + frontier AI tools

πŸ›ƒ Visa sponsorship

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’