Added
11 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

datadog terraform aws grafana prometheus
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now โ†’

๐Ÿ“‹ Description

  • Define and drive SRE practices from the ground up - SLOs, SLIs, error budgets
  • Drive reliability and operational excellence of the Forward SaaS platform
  • Build and maintain observability infrastructure - logging, metrics, tracing, alerting
  • Lead incident response: on-call rotations, runbooks, post-mortems
  • Partner with engineering to embed reliability into the SDLC - capacity planning, load testing
  • Help define and grow the SRE team as a foundational hire with leadership path

๐ŸŽฏ Requirements

  • 6+ years in SRE/DevOps/infra for SaaS or cloud
  • Proven experience building or maturing an SRE function
  • Strong networking fundamentals: TCP/IP, DNS, routing, load balancing
  • Hands-on Kubernetes and container orchestration in production
  • Observability tooling - Prometheus, Grafana, Datadog
  • Cloud platforms (AWS, GCP, Azure) with IaC (Terraform, Ansible)
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’