Senior Site Reliability Engineer

Added
27 days ago
Type
Full time
Salary
Salary not provided

Related skills

terraform grafana prometheus python kubernetes

πŸ“‹ Description

  • Design, build, and evolve scalable cloud infra on GCP
  • Develop tooling to improve autonomy and efficiency
  • Advance observability across metrics, logging, tracing
  • Improve cloud efficiency via cost visibility and governance
  • Champion SRE practices: SLOs, SLIs, error budgets, DORA
  • Lead incident response and post-incident reviews

🎯 Requirements

  • 3+ years in SRE or production-focused roles
  • Strong GCP expertise including cost optimization
  • Experience managing Kubernetes in production
  • IaC with Terraform, Deployment Manager, or similar
  • Strong scripting/automation: Python, Bash, Go
  • Observability tooling: Prometheus, Grafana, OpenTelemetry

🎁 Benefits

  • Fully remote work with eligible locations
  • Ownership of infrastructure and tooling
  • Automation and self-service tooling to boost productivity
  • Exposure to GCP, Kubernetes, Terraform, Prometheus, Grafana, OpenTelemetry
  • Opportunity to contribute to reliability practices and DORA metrics
  • Salary discussed based on experience and location
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’