Lead Site Reliability Engineer - Imunify Reliability Platform

Added
4 days ago
Type
Full time
Salary
Salary not provided

Related skills

rust grafana prometheus python go

πŸ“‹ Description

  • Define SLIs, SLOs, and ownership for ~70 components with squads and senior engineers.
  • Develop a reliability taxonomy covering availability, latency, telemetry health, and configuration
  • Ensure indicators are independently measurable and cannot be defeated by the failure they detect.
  • Design telemetry collection for customer-hosted agents and cloud services with privacy and data
  • Extend instrumentation across Python, Go, and Rust components.
  • Consolidate dashboards and reporting into a reliable observability platform.

🎯 Requirements

  • Substantial production engineering or SRE experience with SLO framework design and implementation.
  • Strong Python skills; able to modify Go or Rust code for instrumentation.
  • Hands-on experience with Prometheus/OpenMetrics, Grafana, Alertmanager, and ClickHouse or similar
  • Experience debugging distributed systems beyond Kubernetes-centric ops.
  • Practical experience with Ansible, GitLab CI, Jenkins, or similar CI/CD tooling.
  • Strong telemetry understanding for non-scrapable systems, including privacy considerations.

🎁 Benefits

  • Fully remote work with flexible hours (Worldwide).
  • 24 vacation days/year + 10 paid holidays.
  • Unlimited sick leave.
  • Private medical insurance compensation and co-working space reimbursement.
  • Gym/sports reimbursement and professional development opportunities.
  • Remote-first, asynchronous work across time zones; opportunity to shape SRE function.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’