Lead Site Reliability Engineer - Imunify Reliability Platform

Added
4 days ago
Type
Full time
Salary
Salary not provided

Related skills

sre site reliability grafana prometheus ci/cd

πŸ“‹ Description

  • Define SLIs for ~70 components with ownership and SLOs.
  • Develop reliability taxonomy for availability, latency, telemetry.
  • Ensure indicators are measurable and cannot be disabled.
  • Design telemetry pipeline for agents and cloud services.
  • Extend instrumentation in Python, Go, Rust.
  • Consolidate dashboards into a unified observability platform.

🎯 Requirements

  • Substantial prod engineering or SRE experience with an SLO framework.
  • Strong Python skills; read/modify Go or Rust for instrumentation.
  • Experience with Prometheus/OpenMetrics, Grafana, Alertmanager, ClickHouse or similar.
  • Debug distributed systems on bare metal/long-lived hosts beyond Kubernetes.
  • Production CI/CD tooling: Ansible, GitLab CI, Jenkins, or similar.
  • Telemetry for systems not directly scraped: push-based collection, sampling, privacy.

🎁 Benefits

  • Fully remote work with flexible hours worldwide.
  • 24 paid vacation days per year.
  • 10 paid national holidays.
  • Unlimited sick leave.
  • Private medical insurance reimbursement.
  • Co-working space reimbursement.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’