Lead Site Reliability Engineer - Imunify Reliability Platform

Added
7 days ago
Type
Full time
Salary
Salary not provided

Related skills

rust grafana prometheus python kubernetes

πŸ“‹ Description

  • Define SLIs for ~70 components with ownership, SLOs, and error budgets.
  • Build reliability taxonomy for availability, telemetry, and config convergence.
  • Ensure reliability indicators are independently measurable.
  • Design telemetry pipelines for agents and cloud services.
  • Extend instrumentation across Python, Go, Rust code.
  • Consolidate dashboards and retire non-value tooling.

🎯 Requirements

  • Significant production engineering or SRE exp with SLO framework work.
  • Strong Python; ability to read/modify Go or Rust for instrumentation.
  • Experience with time-series telemetry: Prometheus/OpenMetrics, Grafana, Alertmanager, ClickHouse.
  • Debugging distributed systems beyond Kubernetes.
  • Experience with Ansible, GitLab CI, Jenkins, or similar.
  • Telemetry for non-scrapable/masked systems; privacy considerations.

🎁 Benefits

  • Fully remote work with flexible hours (worldwide).
  • 30+ days vacation equivalent (24 days) per year.
  • 10 paid national holidays.
  • Unlimited sick leave.
  • Private medical insurance contribution.
  • Co-working space reimbursement.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’