Senior Site Reliability Engineer - Managed Kubernetes

Added
4 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

gitops helm prometheus python kubernetes

πŸ“‹ Description

  • Operate and maintain bare-metal Kubernetes clusters, scaling up to thousands of nodes
  • Handle cluster degradation, recovery, resizing, and incident response using fleet management tools
  • Participate in a well-managed on-call rotation for critical incidents
  • Assist customers with Kubernetes questions, workload integration, storage, and authentication
  • Work closely with our HPC Ops and Datacenter Ops teams for low-level or cross-functional issues
  • Use Python and Golang to create tooling and automate the validation of platform quality
  • Design, build, and maintain scalable control plane services, operators, and custom controllers for
  • Develop automation for cluster lifecycle management: provisioning, upgrades, patching, and deletion
  • Define and implement SLOs and SLIs for Kubernetes services, workloads, and platform reliability
  • 6+ years of experience in a SRE, operations engineer, or similar role, with a deep knowledge of
  • Strong programming skills in Go and Python; experience with GitOps (e.g., ArgoCD), Helm, and
  • Proven experience operating Kubernetes clusters in production environments (on-prem, EKS, GKE, or
  • Can work either independently with limited direction or as part of a team
  • Can work with customers during incidents either via tickets, live messaging, or as part of a larger
  • Familiarity with observability tools like Prometheus, Grafana, FluentBit, and CI/CD pipelines

Nice To Have

  • Deep Kubernetes expertise: CRDs, CSI, CNI, Kubernetes Operator Coding experience
  • Exposure to HPC clusters, AI/ML workloads, or large-scale GPU clusters
  • Hybrid or multi-cloud Kubernetes environment experience
  • Contributions to CNCF projects or Kubernetes SIGs

Why Join Us

  • Work on cutting-edge Managed Kubernetes platforms for AI/ML workloads
  • Influence the platform roadmap and help shape operations and reliability best practices
  • Collaborate with a highly skilled engineer
  • Opportunity to mentor and grow within a fast-growing, technology-driven environment
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’