Lead Site Reliability Engineer - Imunify Reliability Platform

Added
1 day ago
Type
Full time
Salary
Salary not provided

Related skills

rust ansible grafana prometheus python

πŸ“‹ Description

  • Define SLIs/SLOs and error budgets for ~70 components.
  • Build telemetry for customer-hosted agents and cloud services.
  • Extend instrumentation in Python, Go, Rust.
  • Consolidate dashboards and retire ineffective tooling.
  • Implement SRE practices: alerting, escalation, on-call ownership.
  • Lead reliability culture and measurable outcomes across teams.

🎯 Requirements

  • Substantial production engineering or SRE experience with SLO framework design.
  • Strong Python; ability to read/modify Go or Rust for instrumentation.
  • Experience with time-series data at scale: Prometheus/OpenMetrics, Grafana, Alertmanager
  • Experience debugging distributed systems beyond Kubernetes-centric view.
  • Hands-on with CI/CD tooling (Ansible, GitLab CI, Jenkins).
  • Telemetry for push-based collection, privacy, and data quality on customer-managed infra.

🎁 Benefits

  • Fully remote work with flexible hours (worldwide).
  • Competitive compensation and private medical insurance support.
  • Co-working space and gym reimbursements.
  • Professional development opportunities and mentorship.
  • Remote-first, asynchronous environment across time zones.
  • Opportunity to shape SRE function and operational culture.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’