Lead Site Reliability Engineer - Imunify Reliability Platform

Added
6 days ago
Type
Full time
Salary
Salary not provided

Related skills

rust grafana prometheus python go

πŸ“‹ Description

  • Define SLIs for ~70 components with ownership, SLOs, and error budgets.
  • Build telemetry pipeline for customer-hosted agents and cloud services.
  • Extend instrumentation across Python, Go, and Rust components.
  • Consolidate dashboards and reporting into a reliable observability platform.
  • Implement SRE-focused alerting with clear escalation and runbooks.
  • Establish blameless postmortems and cross-team reliability practices.

🎯 Requirements

  • Production engineering or SRE experience with SLO framework design.
  • Strong Python; ability to read/modify Go or Rust.
  • Hands-on time-series telemetry: Prometheus/OpenMetrics, Grafana, Alertmanager, ClickHouse or
  • Experience debugging distributed systems on bare metal/long-lived hosts (not just Kubernetes).
  • CI/CD tooling experience (Ansible, GitLab CI, Jenkins).
  • Telemetry for non-scrapable systems; privacy and partial reporting considerations.

🎁 Benefits

  • Fully remote work with flexible hours (worldwide).
  • 24 vacation days, 10 holidays, unlimited sick leave.
  • Private medical insurance contribution, co-working and gym reimbursements.
  • Professional development through challenging projects and mentoring.
  • Option to define SRE function and operational culture from the ground up.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’