Senior Site Reliability Engineer

Added
4 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

terraform aws grafana prometheus kubernetes

πŸ“‹ Description

  • Operate and maintain all Develocity instances and supporting services.
  • Participate in follow-the-sun on-call rotation; own incident response and troubleshooting.
  • Drive automation across deployment, upgrades, monitoring, self-healing, and recovery.
  • Build and maintain observability for all managed services (logging, metrics, tracing, and alerting).
  • Work with engineering teams to build reliability into features from the start.
  • Run incident retrospectives, and own disaster recovery, backups, and business continuity.

🎯 Requirements

  • 5+ years in SRE/DevOps operating production services at scale
  • Strong Kubernetes experience in production environments
  • Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2)
  • Proficiency with observability tools (Prometheus, Grafana) and IaC (Terraform)
  • Track record of incident management and response; 24/7 on-call rotations
  • Knowledge of SRE best practices (SLAs, SLOs)

🎁 Benefits

  • A ground-floor role in a new SRE team
  • Real ownership of production systems used by engineers at companies you've heard of
  • Direct interaction with customers during incidents
  • A culture that rewards automation over heroics
  • Work from home in a remote-first environment
  • Competitive salaries and equity grants
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’