VP Engineering - Infrastructure & SRE

Added
1 hour ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

terraform aws postgresql grafana prometheus
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now β†’

πŸ“‹ Description

  • Define and execute the infrastructure vision, strategy, and multi-year roadmap to support platform growth, operational resilience, and compliance requirements.
  • Lead, mentor, and scale infrastructure, platform, and SRE teams by developing talent, improving processes, and fostering a culture of ownership and operational excellence.
  • Own platform reliability by establishing service level objectives (SLOs), error budgets, incident management practices, and measurable reliability improvements.
  • Participate in and oversee the on-call rotation, acting as a senior incident leader during critical events by driving investigation, communication, mitigation, and long-term resolution.
  • Direct cloud infrastructure strategy, including AWS architecture, account structures, networking, identity management, multi-region capabilities, and cost optimization initiatives.
  • Own Kubernetes platform strategy, including architecture, upgrades, workload management, scaling, developer experience, and operational best practices.

🎯 Requirements

  • 15+ years of experience across infrastructure, platform engineering, site reliability, software development, or related engineering disciplines.
  • 5+ years of experience leading engineering teams, including hiring, coaching, organizational design, and performance development.
  • Extensive hands-on expertise designing and operating production AWS environments, including compute, networking, IAM, multi-account architectures, and cloud cost management.
  • Deep production experience with Kubernetes, including cluster operations, workload architecture, scaling strategies, and platform reliability.
  • Strong expertise with relational databases at scale, particularly Aurora RDS, MySQL, and/or PostgreSQL, including high availability, replication, performance tuning, and recovery strategies.
  • Proven experience owning disaster recovery and business continuity programs with measurable recovery objectives and tested failover processes.

🎁 Benefits

  • Competitive total compensation package with an estimated range of $400,000 - $600,000 USD annually.
  • Equity opportunities and retirement benefits including a 401(k) match.
  • Comprehensive medical, dental, and vision insurance coverage.
  • Life, short-term disability, and long-term disability insurance.
  • Unlimited paid time off, volunteer hours, and sabbatical opportunities.
  • Opportunity to work remotely from the United States.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’