Lead Site Reliability Engineer

Added
10 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

terraform aws prometheus python kubernetes

πŸ“‹ Description

  • Define automation strategy for infrastructure tooling to reduce incidents.
  • Own design and reliability of core platform apps; mentor teams.
  • Architect logging platform, balancing availability, retention, cost.
  • Establish capacity planning and performance management frameworks.
  • Lead cross-functional reliability initiatives with SRE and service teams.
  • Demonstrate high autonomy in addressing systemic platform weaknesses.

🎯 Requirements

  • Proven track record in SRE/Software Eng building scalable, reliable services.
  • Deep expertise with distributed systems: Pulsar, Kafka, Loki, ScyllaDB/Cassandra.
  • Automation strategies and performance analysis; mentoring on diagnostics and resolution.
  • 6+ years hands-on in SRE or Software Eng; multi-cloud AWS & GCP.
  • Observability platforms, SLOs; Prometheus, Thanos, Grafana Loki, Tempo.
  • Terraform and Chef IaC, Advanced Kubernetes (EKS/GKE), multi-tenancy.

🎁 Benefits

  • Medical, financial, and other benefits.
  • Flexible, inclusive culture; equal opportunity employer.
  • Opportunity to impact a platform serving billions of requests daily.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’