Lead Site Reliability Engineer - Imunify Reliability Platform

Added
1 day ago
Type
Full time
Salary
Salary not provided

Related skills

rust ansible grafana prometheus python

πŸ“‹ Description

  • Define SLIs/SLOs for ~70 components with ownership and error budgets.
  • Build telemetry for customer-hosted agents and cloud services.
  • Extend instrumentation across Python, Go, and Rust.
  • Consolidate dashboards into a reliable observability platform.
  • Implement SLO-driven alerting with multi-window burn rates.

🎯 Requirements

  • Production engineering or SRE exp; define an SLO framework.
  • Strong Python; read/modify Go or Rust for instrumentation.
  • Time-series/event telemetry at scale; Prometheus/OpenMetrics, Grafana, Alertmanager, ClickHouse.
  • Debug distributed systems beyond Kubernetes focus.
  • Experience with Ansible, GitLab CI, Jenkins, or similar.
  • Telemetry for systems with push-based collection and privacy concerns.

🎁 Benefits

  • Fully remote work with global hours.
  • Paid vacation, holidays, sick leave.
  • Private medical insurance contribution; co-working space reimbursement.
  • Gym/sports reimbursement; pro development opportunities.
  • Opportunity to shape SRE function and culture.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’