Senior Site Reliability Engineer

Added
16 minutes ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

datadog terraform helm grafana prometheus

πŸ“‹ Description

  • Own Data Replication infra: Kubernetes, CI/CD, secrets, networking, cloud config.
  • Partner with product engineers to reliably integrate features with infrastructure.
  • Maintain observability, alerting, and anomaly detection with AI/LLM focus.
  • Maintain AI-augmented release tooling: canaries, progressive rollouts, rollback.
  • Build self-serve tooling, write runbooks, and coach engineers to own more of their stack.

🎯 Requirements

  • 7+ years in infrastructure, platform engineering, SRE, or DevOps.
  • Hands-on ownership of Kubernetes, Helm, and Terraform in production environments.
  • Experience with observability stacks (Prometheus, Grafana, Datadog) and on-call operations.
  • Experience with CI/CD pipeline ownership and developer tooling.
  • Ability and willingness to read backend code to understand and instrument failures.
  • Fluency with AI tools - LLMs and agentic frameworks to automate and reduce toil.
  • Startup-ready mindset: comfortable with ambiguity, moving fast, owning problems end-to-end.

🎁 Benefits

  • Flexible PTO with at least 25 days off annually.
  • 16 weeks fully paid parental leave for all parents.
  • Comprehensive medical, dental, and vision coverage for employees and dependents.
  • 401(k) retirement plan.
  • Professional development budget, conference sponsorship, and book reimbursement.
  • Commuter benefits and monthly internet reimbursement.
  • Breakfast and lunch in our San Francisco office.
  • Collaborative, in-person culture focused on learning, growth, and impact.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’