Senior Site Reliability Engineer (Arlington, VA) - Relocation Provided

Added
less than a minute ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

sre ansible terraform aws prometheus

πŸ“‹ Description

  • Design, implement, and manage monitoring, logging, and alerting stack (e.g., Prometheus, Loki
  • Define, measure, and own alerting that feeds into Service Level Indicators (SLIs) and Service Level
  • Act as incident responder and potentially incident commander during critical incidents, leading
  • Partner with platform engineers to design, build, and manage secure, resilient Kubernetes clusters
  • Embed security and compliance controls (RMF, STIGs) into automation
  • Proactively identify and eliminate operational toil by building automation

🎯 Requirements

  • Active Top Secret clearance
  • 5+ years in Platform, DevOps, or Site Reliability Engineering with an infrastructure and operations
  • Proven partner to DevOps/Platform and application teams; collaborates well across functions and
  • Deep understanding of incident response processes with experience conducting thorough root cause
  • Infrastructure as Code: Terraform (or CloudFormation), Ansible
  • Kubernetes design, deployment, and operations

🎁 Benefits

  • $180K – $220K salary
  • Offers Equity

🚚 Relocation support

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’