Staff Site Reliability Engineer

Added
9 days ago
Type
Full time
Salary
Salary not provided

Related skills

gitops terraform helm python kubernetes

๐Ÿ“‹ Description

  • Design, build, and operate large-scale cloud infrastructure and production services.
  • Participate in an on-call rotation supporting highly available customer-facing systems.
  • Lead incident response and drive post-incident reviews for systemic improvements.
  • Define, measure, and improve SLIs, SLOs, and error budgets.
  • Partner with engineering teams to improve availability, scalability, and resilience.
  • Continuously improve observability through metrics, logging, tracing, dashboards.

๐ŸŽฏ Requirements

  • AWS and/or GCP production services experience.
  • Kubernetes production expertise; troubleshoot networking, storage, scaling.
  • Terraform and Helm IaC; GitOps practices.
  • Strong Go and/or Python software engineering.
  • Experience building automation and internal platforms.
  • Observability/telemetry; PostgreSQL, Redis, OpenSearch.

๐ŸŽ Benefits

  • Immersive in-person onboarding.
  • Well-being programs and benefits.
  • Social impact initiatives.
  • Talent development and community connections.
  • Inclusive culture and equal opportunity employer.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’