Lead Site Reliability Engineer - Imunify Reliability Platform

Added
4 days ago
Type
Full time
Salary
Salary not provided

Related skills

sre site reliability grafana prometheus clickhouse

๐Ÿ“‹ Description

  • Define SLIs/SLOs for ~70 product components with ownership, measurement, and error budgets.
  • Build a reliability platform and telemetry pipeline for cloud services and customer-hosted agents.
  • Extend instrumentation across Python, Go, and Rust; consolidate dashboards and observability
  • Implement SLO-driven alerting with multi-window burn-rate and clear classifications.
  • Establish escalation models, on-call practices, and blameless postmortems across time zones.
  • Lead discussions to shape reliability culture and measurable outcomes.

๐ŸŽฏ Requirements

  • Substantial production engineering or SRE experience with an SLO framework.
  • Strong Python skills; ability to read/modify Go or Rust for instrumentation.
  • Hands-on experience with Prometheus/OpenMetrics, Grafana, Alertmanager, ClickHouse (or similar).
  • Experience debugging distributed systems beyond Kubernetes-centric operations.
  • Practical experience with CI/CD tooling (Ansible, GitLab CI, Jenkins).
  • Understanding of push-based telemetry, privacy, and data quality for customer-managed infra.

๐ŸŽ Benefits

  • Fully remote work with flexible hours, globally.
  • 28 days total vacation, 10 holidays, unlimited sick leave.
  • Private medical insurance contribution, coworking space, gym reimbursements.
  • Professional development and opportunities to shape an SRE function.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’