Lead Site Reliability Engineer - Imunify Reliability Platform

Added
1 day ago
Type
Full time
Salary
Salary not provided

Related skills

rust grafana prometheus python kubernetes

πŸ“‹ Description

  • Define SLIs for ~70 components with ownership, measurement, SLOs, and error budgets
  • Develop a reliability taxonomy for availability, latency, telemetry health
  • Ensure indicators are independently measurable
  • Design telemetry pipelines for customer-hosted agents and cloud services
  • Extend instrumentation across Python, Go, and Rust
  • Consolidate dashboards and establish observable platform

🎯 Requirements

  • Substantial production engineering or SRE experience with SLO framework
  • Strong Python skills; ability to modify Go/Rust for instrumentation
  • Hands-on time-series/event telemetry at scale (Prometheus/OpenMetrics, Grafana, Alertmanager
  • Experience debugging distributed systems beyond Kubernetes-centric setups
  • Practical experience with CI/CD tooling (Ansible, GitLab CI, Jenkins)
  • Understanding telemetry for non-scraped systems, including privacy considerations

🎁 Benefits

  • Fully remote work with flexible hours
  • 24 vacation days, 10 holidays, unlimited sick leave
  • Private medical insurance support, co-working space reimbursement
  • Gym/sports reimbursement, professional development opportunities
  • Opportunity to define SRE function and reliability culture
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’