Senior Cloud Site Reliability Engineer

Added
7 days ago
Type
Full time
Salary
Salary not provided

Related skills

gitops jenkins terraform github actions helm

๐Ÿ“‹ Description

  • Design and run scalable, reliable systems across multi-clouds (AWS)
  • Drive uptime, latency, and service health with SLOs/SLIs/SLAs
  • Own observability stack and incident readiness (Prometheus, Grafana, OpenTelemetry)
  • Automate infrastructure with Terraform, Helm, Kubernetes
  • Lead incident response, RCA, and partner with product teams for readiness
  • Mentor junior SREs and contribute to reliability roadmaps

๐ŸŽฏ Requirements

  • 5+ years with Kubernetes, EKS, ECS, and containerized workloads
  • AWS services: EC2, Lambda, IAM, RDS, S3, ALB/NLB, VPC, PrivateLink
  • Terraform, Helm, Jenkins, GitHub Actions, and GitOps (ArgoCD/Flux)
  • Observability: Prometheus, Grafana, OpenTelemetry, and distributed monitoring
  • Linux, networking fundamentals, and performance tuning
  • Python, Go, or Shell scripting for automation

๐ŸŽ Benefits

  • NICE-FLEX hybrid model: 2 days in office, 3 remote
  • Global company with internal mobility across roles and locations
  • Career growth opportunities and learning culture
  • Equal opportunity employer
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’