Specialist Cloud Site Reliability Engineer

Added
7 days ago
Type
Full time
Salary
Salary not provided

Related skills

jenkins terraform github actions helm aws

๐Ÿ“‹ Description

  • Design and implement scalable, reliable systems across multi-cloud (AWS/EKS/Lambda)
  • Improve uptime, latency, and service health with SLOs/SLIs/SLAs
  • Build IaC and automation with Terraform, Helm, Kubernetes
  • Enhance CI/CD pipelines with Jenkins, GitHub Actions; safe rollouts and rollbacks
  • Own observability stack: Prometheus, Grafana, Loki, Tempo, OpenTelemetry
  • Lead incident response, RCA, postmortems; prep for production releases

๐ŸŽฏ Requirements

  • 8+ years with Kubernetes, EKS, ECS and containerized prod workloads
  • AWS: EC2, Lambda, IAM, RDS, S3, VPC, ALB/NLB
  • Terraform, Helm, Jenkins, GitHub Actions, GitOps (ArgoCD/Flux)
  • Observability: metrics, logs, traces, distributed monitoring
  • Prometheus, Grafana, Loki, Tempo, OpenTelemetry
  • Linux, networking, and system performance tuning
  • Python/Go/Shell scripting for automation
  • Incident response, RCA, on-call operations

๐ŸŽ Benefits

  • NICE-FLEX hybrid model: 2 days in office, 3 days remote
  • Global opportunities across roles and locations
  • Equal opportunity employer
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’