Platform Site Reliability Engineer

Added
3 hours ago
Type
Full time
Salary
Salary not provided

Related skills

ansible linux ubuntu grafana prometheus

๐Ÿ“‹ Description

  • Deploy and manage Kubernetes clusters at scale for AI workloads across infrastructure.
  • Develop Kubernetes manifests and Operators for deployments, networking, storage, security.
  • Optimize Linux configs (kernel, drivers, FS) for the orchestration layer.
  • Build automation scripts and IaC for platform lifecycle; aid incident resolution.
  • Maintain observability stack: Prometheus, Grafana, and custom monitoring.
  • Operate 24x7 production environments with on-call; contribute to postmortems.

๐ŸŽฏ Requirements

  • 5+ years in globally scaled, 24/7 SRE or equivalent.
  • 3+ years running, deploying and optimizing Kubernetes.
  • Expert-level Linux administration, especially Ubuntu.
  • Proficiency in system tuning, disk I/O, and hardware-level tweaks.
  • Strong networking fundamentals: TCP/IP, DNS, DHCP, VLANs, routing, switching.
  • Strong experience with infrastructure scripting and automation (Bash, Python, Ansible).

๐ŸŽ Benefits

  • 25 days of annual leave
  • Private medical insurance via Bupa
  • Cycle to Work Scheme
  • Gympass subscription
  • Participation in the company shares program
  • Enhanced parental pay & leave
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’