Principal Platform Engineer (High Availability & Disaster Recovery)

Added
13 days ago
Type
Full time
Salary
Salary not provided

Related skills

azure ansible terraform aws kubernetes

๐Ÿ“‹ Description

  • Own the platform's HA, resilience, and DR lifecycle design and execution.
  • Design and deliver solutions for availability, resiliency, and disaster response.
  • Set guardrails for fault-tolerance and lead multi-region DR planning across cloud and on-prem.
  • Build automated failover mechanisms and champion chaos engineering practices.
  • Mitigate single points of failure and coordinate across global teams.
  • Drive end-to-end HA architecture and roadmap for the platform.

๐ŸŽฏ Requirements

  • 8+ years of HA engineering across cloud and on-prem.
  • Experience with AWS, GCP, and Azure across multi-region environments.
  • Hands-on IaC with Terraform and/or Ansible.
  • Expertise in active-active clustering, global load balancing, and real-time DB replication.
  • Kubernetes experience; Linux and containerization.
  • Incident management, post-mortems, and disaster recovery testing.
  • Strong cross-functional communication and strategic execution.

๐ŸŽ Benefits

  • Winning culture with opportunities to learn, grow, and be rewarded.
  • Inclusive, diverse workplace with DEIB programs.
  • Global exposure with Fortune 50 clients.
  • Competitive compensation and comprehensive benefits.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’