Senior Site Reliability Engineer (Golang / Kubernetes)

Added
17 days ago
Type
Full time
Salary
Salary not provided

Related skills

prometheus python kubernetes go slo

๐Ÿ“‹ Description

  • Define SLIs and SLOs based on the signals available across the platform โ€” Kubernetes, bare-metal
  • Design and build the API that exposes SLIs and reliability state to Platform Administrators and
  • Establish alerting and error-budget practices that maximize signal and minimize noise.
  • Partner with infrastructure, storage, and networking teams to ensure the right signals are
  • Diagnose reliability and performance issues across the observability stack and drive their

๐ŸŽฏ Requirements

  • 5+ years in SRE, platform reliability, or a closely related software/infrastructure role.
  • Strong software engineering skills (e.g., Go or Python) with experience building and operating APIs
  • Demonstrated experience defining SLIs/SLOs and error budgets for real production systems.
  • Hands-on experience with observability tooling โ€” metrics, logging, and tracing (e.g.
  • Solid understanding of Kubernetes and the signals it and its workloads emit.
  • Strong written and verbal communication with technical audiences.

๐ŸŽ Benefits

  • Work with an established Silicon Valley leader in the cloud infrastructure industry
  • Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and
  • Be a part of cutting-edge, open-source innovation
  • Thrive in the high-energy environment of a young company where openness, collaboration
  • Professional development and training
  • Attend conferences and working groups
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’