Senior Site Reliability Engineer

Added
1 day ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

terraform github actions aws grafana prometheus

πŸ“‹ Description

  • Lead discovery, design, and delivery of reliability initiatives
  • Shape platform architecture and tooling priorities for reliability
  • Define SLOs, SLIs, error budgets, alerts, observability
  • Use metrics to identify systemic reliability issues
  • Resolve cross-team requests, automate and document recurring problems
  • Operate and scale production Kubernetes and container infra

🎯 Requirements

  • Experience in SRE/DevOps/Platform Engineering
  • Production Kubernetes with Docker and container ecosystem
  • Production cloud infra using AWS or similar
  • Terraform and IaC expertise
  • Experience with SLOs/SLIs, alerting, incident mgmt
  • Observability with OpenTelemetry, Grafana, Prometheus

🎁 Benefits

  • 100% remote work, anywhere
  • Async-friendly hours
  • Flexible paid time off
  • 16 weeks parental leave
  • Budget for coworking spaces, learning, wellness
  • Mental health support services
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’