Software Engineer III, Site Reliability

Added
8 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

datadog terraform python kubernetes go
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now โ†’

๐Ÿ“‹ Description

  • Own and evolve SLI/SLO and error budgets to guide priorities.
  • Lead incident response, postmortems, and systemic fixes.
  • Build and maintain observability across metrics, logs, traces using Datadog.
  • Design and operate resilient, scalable infrastructure with Terraform.
  • Manage production Kubernetes and container workloads; plan capacity and optimize cloud costs.
  • Own CI/CD pipelines and safe deployment strategies (canary, rollout, rollback).

๐ŸŽฏ Requirements

  • 5+ years in site reliability, platform, or infrastructure engineering with production ownership.
  • Strong automation skills in Go, Python, or TypeScript; building tooling.
  • Hands-on AWS, Kubernetes, and Terraform experience.
  • Proven track record leading incident response and SLO-driven reliability.
  • Experience with observability tools like Datadog.
  • Security in CI/CD: SAST/DAST/SCA tooling and policy-as-code.

๐ŸŽ Benefits

  • Flexible time-off policy with Responsible Time Off.
  • Healthcare, dental, and vision benefits.
  • 401(k) plan with employer match.
  • Monthly wellness and tech allowances.
  • Mentorship program.
  • Paid parental leave and fertility benefits.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’