Lead Site Reliability Engineer - Imunify Reliability Platform

Added
5 days ago
Type
Full time
Salary
Salary not provided

Related skills

rust ansible grafana prometheus python

πŸ“‹ Description

  • Lead SRE for Imunify Reliability Platform in a remote setting
  • Define SLIs/SLOs and error budgets across ~70 components
  • Build a reliability taxonomy for availability, latency, telemetry, and security
  • Architect telemetry collection for cloud services and agent-based workloads
  • Instrument Python, Go, and Rust components for reliability gains
  • Unify dashboards and reporting into a streamlined observability platform

🎯 Requirements

  • Production SRE or engineering experience with SLO framework creation
  • Strong Python; reading/modifying Go or Rust for instrumentation
  • Experience with time-series telemetry: Prometheus/OpenMetrics, Grafana, Alertmanager
  • Hands-on with scalable telemetry in colum nar/high-cardinality stores (e.g., ClickHouse)
  • Debugging distributed systems beyond Kubernetes-centric scope
  • CM/CI tooling: Ansible, GitLab CI, Jenkins

🎁 Benefits

  • Fully remote, flexible hours, work from anywhere
  • 24 vacation days/year
  • 10 paid national holidays
  • Unlimited sick leave
  • Private medical insurance contribution
  • Co-working space reimbursement
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’