Senior Site Reliability Engineer

Added
15 days ago
Type
Full time
Salary
Salary not provided

Related skills

aws python kubernetes go google cloud platform

๐Ÿ“‹ Description

  • Operate and improve the multi-tenant Kubernetes infrastructure that runs customer workloads
  • Build reliability into services and infrastructure; ensure resilience and self-healing
  • Define metrics to detect incidents and measure service health
  • Participate in 24/7 on-call rotation to resolve platform infra issues
  • Mentor early-career SREs and contribute to ops practices

๐ŸŽฏ Requirements

  • Strong background in software development and operating distributed systems
  • 6+ years building and operating distributed systems; proficient in Python or Go
  • Production Kubernetes experience: scheduling, networking, and node-level issues
  • Cloud platforms: AWS, Google Cloud Platform (GCP), or Azure
  • Linux internals and networking: TCP/IP, DNS, TLS, routing
  • Customer-focused mindset and strong verbal/written communication skills

๐ŸŽ Benefits

  • Leadership commitment and inclusive culture
  • Employee affinity groups
  • Fertility assistance and generous parental leave policy
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’