Staff Site Reliability Engineer-Observability

Added
24 hours ago
Type
Full time
Salary
Salary not provided

Related skills

ansible terraform aws grafana prometheus

📋 Description

  • Design and improve observability using Prometheus, Grafana, OpenTelemetry, Elastic.
  • Lead reliability and performance for cloud-native and hybrid infra (AWS/K8s).
  • Manage patching, health, and SLAs across on-prem and cloud deployments.
  • Operate and optimize Kubernetes and GitOps, ArgoCD, multi-tenant envs.
  • Build automation and AI-assisted tooling to speed incident response.
  • Partner with engineering to define metrics, alerts, and readiness; mentor teams.

🎯 Requirements

  • 8+ years in designing and operating large AWS cloud infra.
  • 5+ years with Kubernetes platforms (EKS/AKS/GKE, Fargate).
  • 4+ years Linux admin in production.
  • Go, Python, or Ruby for automation.
  • Terraform, Ansible, AWS CDK or similar IaC.
  • Observability, monitoring, logging, tracing basics.
  • CI/CD pipelines, Git, automated testing.
  • MySQL or PostgreSQL, networking fundamentals, security.

🎁 Benefits

  • Fully remote in India.
  • Competitive compensation with bonuses.
  • Equity and stock purchase where applicable.
  • Health and wellness benefits for dependents.
  • Retirement savings and local statutory benefits.
  • Generous PTO, holidays, parental leave.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →