Senior Site Reliability Engineer

Added
12 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

postgresql grafana prometheus kubernetes go

📋 Description

  • Mentor observability best practices, SLIs/SLOs, and reliability culture across teams.
  • Contribute to triage and remediation processes as a player/coach.
  • Lead incident response and debug production issues across the stack.
  • Design and maintain core infra and tooling used by Tulip’s engineering teams.

🎯 Requirements

  • 5+ years with open-source Observability tools (Loki, Grafana, Tempo, Mimir).
  • Hands-on OpenTelemetry instrumentation and Prometheus metrics pipelines at scale.
  • Experience developing AI processes (Claude Skills, Gemini Gems) and iterating on efficacy.
  • Experience with time-series data, ideally PromQL.

🎁 Benefits

  • Direct impact on product and culture
  • Company equity
  • Comprehensive benefits: Health, Dental, Vision, Disability, Life, FSA, 401(k).
  • Flexible work schedule and unlimited vacation policy
  • Virtual company events and happy hours
  • Fitness subsidies
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →