Lead Site Reliability Engineer - Imunify Reliability Platform

Added
3 days ago
Type
Full time
Salary
Salary not provided

Related skills

rust grafana prometheus python go

πŸ“‹ Description

  • Define SLIs, SLOs, and error budgets for ~70 components with squad leads and senior engineers.
  • Develop a reliability taxonomy covering availability, latency, telemetry health, and
  • Ensure indicators are independently measurable and not disabled by the triggering failure.
  • Design telemetry for customer-hosted agents and cloud services; balance push collection, sampling
  • Extend instrumentation across Python, Go, and Rust components.
  • Consolidate dashboards and reporting into a reliable observability platform; retire non-value

🎯 Requirements

  • Substantial production engineering or SRE experience; experience defining and implementing an SLO
  • Strong Python skills; ability to read/modify Go or Rust for instrumentation.
  • Hands-on telemetry experience at scale: Prometheus/OpenMetrics, Grafana, Alertmanager, ClickHouse
  • Experience debugging distributed systems beyond Kubernetes-centric view.
  • Practical experience with production CI/CD tooling (Ansible, GitLab CI, Jenkins, etc.).
  • Understanding telemetry for systems that cannot be fully scraped, including push-based collection

🎁 Benefits

  • Fully remote work with flexible hours, work from anywhere worldwide.
  • 24 paid vacation days per year.
  • 10 paid national holidays.
  • Unlimited sick leave.
  • Private medical insurance contribution.
  • Co-working space reimbursement.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’