Site Reliability Engineer

Added
2 hours ago
Type
Full time
Salary
Salary not provided

Related skills

datadog azure terraform powershell python

πŸ“‹ Description

  • Own the availability and performance of production SaaS apps on Azure across multiple regions.
  • Lead troubleshooting of cloud infra and app issues (AKS pod/node failures, rollbacks, networking
  • Participate in on-call rotation and drive incident response to minimize customer impact.
  • Improve disaster recovery, failover, and incident management across multi-region deployments.
  • Build and maintain automation scripts and monitoring tools to reduce manual toil.
  • Communicate incident status and post-incident summaries to stakeholders.

🎯 Requirements

  • 5+ years in Site Reliability/DevOps or Cloud Admin with production ownership.
  • Hands-on Azure experience (AKS, App Service, Redis, SQL, Service Bus).
  • Monitoring/logging expertise (Datadog, Azure Monitor, ELK; Datadog APM).
  • Networking fundamentals (firewalls, load balancers, VPNs, DNS, routing).
  • Automation/scripting (PowerShell, Python or similar).
  • Cloud DR/backups with geo-redundant/multi-region deployments.

🎁 Benefits

  • Competitive salaries and meaningful bonus program.
  • Excellent benefits including healthcare, retirement matching, life insurance, EAP, time off
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’