Site Reliability Engineer (SRE)

Added
4 hours ago
Type
Contract
Salary
Salary not provided

Related skills

gitops terraform linux kubernetes opentelemetry
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now โ†’

๐Ÿ“‹ Description

  • Operate and enhance Kubernetes platforms across AWS, Azure, and on-premise environments.
  • Lead incident response, problem management, and root cause analysis activities.
  • Manage cluster lifecycle: upgrades, patching, node pools, CNI/CSI, ingress, Rancher.
  • Own observability strategy including dashboards, alerting, monitoring, and SLOs/SLIs.
  • Implement GitOps practices using Fleet and reduce toil through automation.
  • Apply secure API gateway and Web Application Firewall patterns.

๐ŸŽฏ Requirements

  • Deep expertise in Kubernetes, Rancher, GitOps, Linux, and cloud networking.
  • Strong experience operating in hybrid cloud environments across AWS, Azure, and on-premise.
  • Strong automation and scripting skills in Python, Go, Bash, PowerShell, or .NET.
  • Proven experience with Infrastructure as Code using Terraform and Crossplane.
  • Experience implementing and managing observability tooling including Grafana, Prometheus, Jaeger or Tempo, CloudWatch, Loki, and OpenTelemetry.
  • Experience operating within regulated environments including PCI DSS and GDPR.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’