Lead Site Reliability Engineer - Imunify Reliability Platform

Added
4 days ago
Type
Full time
Salary
Salary not provided

Related skills

sre site reliability grafana prometheus python

πŸ“‹ Description

  • Define SLIs/SLOs and error budgets for ~70 components.
  • Build telemetry for customer-hosted agents and cloud services.
  • Extend instrumentation in Python, Go, and Rust.
  • Consolidate dashboards and observability tooling.
  • Implement multi-window burn-rate alerting and runbooks.
  • Establish ownership, escalation, and on-call practices.

🎯 Requirements

  • Extensive production engineering or SRE with SLO framework experience.
  • Strong Python; read/modify Go or Rust for instrumentation.
  • Hands-on time-series telemetry (Prometheus/OpenMetrics, Grafana, Alertmanager).
  • Experience debugging distributed systems beyond Kubernetes.
  • Production config mgmt and CI/CD (Ansible, GitLab CI, Jenkins).
  • Telemetry for push-based, non-scraped systems; privacy concerns.

🎁 Benefits

  • Fully remote with flexible hours.
  • Paid vacation, holidays, and sick leave.
  • Private medical insurance support, space reimbursement.
  • Gym/sports stipend and professional development.
  • Remote-first, async org with global time zones.
  • Opportunity to define SRE function and culture.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’