Staff Site Reliability Engineer, Environment Automation

Added
1 hour ago
Type
Full time
Salary
Salary not provided

Related skills

sre ansible terraform grafana prometheus

๐Ÿ“‹ Description

  • Build & Scale Multi-Tenant Infrastructure: Design and implement automation that provisions and
  • Debug & Resolve Production Issues: Troubleshoot issues across Kubernetes clusters, cloud
  • Automate Operations at Scale: Replace manual workflows with infrastructure-as-code solutions
  • Monitor & Predict Capacity: Build observability systems that detect bottlenecks, predict usage
  • Respond & Lead During Incidents: Lead incident response and postmortem efforts, applying
  • Architect & Collaborate: Influence architectural decisions around automation, scalability, and

๐ŸŽฏ Requirements

  • Production-Scale Experience: Proven ability to operate and troubleshoot production workloads across
  • Terraform & IaC Mastery: Strong hands-on experience with Terraform, including workspace
  • Kubernetes in Production: Skilled at diagnosing deployment failures, interpreting pod logs, and
  • Programming & Code Analysis: Ability to read and debug code in Go and/or Ruby. Familiar with
  • Large Scale Operations Background: Experience supporting infrastructure for many customers or
  • Architecture & Incident Response: Able to reason through complex systems and operational

๐ŸŽ Benefits

  • Benefits to support your health, finances, and well-being
  • Flexible Paid Time Off
  • Team Member Resource Groups
  • Equity Compensation & Employee Stock Purchase Plan
  • Growth and Development Fund
  • Parental Leave
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’