Senior Site Reliability Engineer

Added
5 hours ago
Type
Full time
Salary
Salary not provided

Related skills

datadog azure docker terraform aws

📋 Description

  • Implement and manage cloud-native systems (AWS) with automation.
  • Operate and optimize Kubernetes clusters, pipelines, and service meshes.
  • Design and maintain multi-tenant SaaS infra with reliability and security.
  • Define and maintain SLOs/SLAs and error budgets; address issues proactively.
  • Develop IaC (Terraform) for repeatable provisioning.
  • Build internal tools for provisioning, monitoring, and ops efficiency.
  • Monitor infrastructure and apps to ensure high-quality user experiences.
  • Participate in on-call rotation; incident response and post-incident reviews.
  • Act as Incident Commander during on-call duty and coordinate cross-team responses.
  • Drive incident response processes to reduce MTTD/MTTR.
  • Diagnose and resolve complex issues independently.
  • Collaborate with engineers to embed observability and reliability into service design.
  • Automate runbooks, health checks, and alerts.
  • Support canary deployments, testing, and rollback strategies.
  • Contribute to security practices, compliance automation, and cost optimization.

🎯 Requirements

  • Bachelor’s degree in CS or equivalent practical experience.
  • 5+ years in SRE/Platform Eng with software dev practices.
  • Hands-on with AWS, GCP, Azure and SaaS products.
  • Programming/scripting: Python, Go, Bash; Terraform.
  • Proficiency in Kubernetes, Docker, Helm.
  • Experience with multi-tenant microservices.
  • CI/CD pipelines and tools: Jenkins, GitHub Actions, GitLab CI.
  • Monitoring tools like Datadog.
  • On-call experience and leading post-incident reviews.
  • Production systems operation with urgency and method.
  • Linux, networking, troubleshooting.
  • Networking: TCP/IP, VPN; VPC, Istio; storage (S3/EBS).
  • Zero-downtime deployments: blue/green, canary.
  • SOC 2, ISO27001, HIPAA; FedRAMP a plus.
  • Chaos engineering/resilience testing.
  • Strong problem-solving and collaboration; agile.
  • Self-driven, organized; prioritize effectively.
  • Curious to learn new technologies.
  • Excellent English communication.
  • Open to diverse backgrounds; apply even if not all requirements.

🎁 Benefits

  • Hybrid work environment.
  • Global team with 75+ nationalities.
  • Opportunity to work with cutting-edge technology.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →