Senior Site Reliability Engineer

Added
16 days ago
Type
Full time
Salary
Salary not provided

Related skills

terraform bash grafana prometheus python

πŸ“‹ Description

  • Design, build, and evolve scalable cloud infrastructure on GCP
  • Develop tooling and self-service capabilities to boost engineering autonomy
  • Improve observability across metrics, logging, tracing, and distributed systems
  • Establish cost visibility and governance to optimize cloud usage
  • Champion reliability practices: SLOs, SLIs, error budgets, DORA metrics
  • Lead incident management and participate in on-call rotation

🎯 Requirements

  • 3+ years in Site Reliability Engineering or production-focused SRE roles
  • Strong hands-on with GCP, cost optimization, governance
  • Experience managing Kubernetes clusters in production
  • IaC with Terraform/Deployment Manager or similar tools
  • Strong scripting/automation: Python, Bash, Go
  • Observability: Prometheus, Grafana, OpenTelemetry; logging/tracing

🎁 Benefits

  • Fully remote work with flexible locations
  • Ownership over infrastructure and reliability practices
  • Automation and self-service tooling to boost productivity
  • Exposure to GCP, Kubernetes, Terraform, Prometheus, Grafana, OpenTelemetry
  • Opportunity to contribute to SLOs/SLIs, incident management, and DORA metrics
  • Salary discussed during interview, based on experience and location
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’