Senior Software Engineer, Infrastructure Engineering

Added
13 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

aws grafana prometheus kubernetes gcp

πŸ“‹ Description

  • Lead incident response and PIRs; drive RCA and long-term fixes.
  • Own incident playbooks; ensure preparedness for diverse failure scenarios.
  • Maintain observability with Prometheus and Grafana; detect bottlenecks early.
  • Automate incident detection and recovery; minimize manual work.
  • Design scalable core services; plan for disaster recovery.
  • Build CI/CD pipelines and document hardware automation and lifecycle.

🎯 Requirements

  • 7+ years in cloud operations, SRE, or related roles.
  • Proficiency in Go; experience deploying containerized apps on Kubernetes.
  • Deep experience with Prometheus and Grafana; observability focus.
  • Cloud platforms: Kubernetes, AWS, and GCP.
  • Familiar with ITIL/SRE incident management; on-call experience.
  • Excellent documentation; strong analytical and problem-solving skills.

🎁 Benefits

  • Medical, dental, and vision insurance - 100% paid for employee.
  • 401(k) with generous employer match.
  • Flexible PTO.
  • Paid Parental Leave.
  • Flexible Spending Account.
  • Catered lunch daily in offices and data centers.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’