Lead Site Reliability Engineer - Imunify Reliability Platform

Added
4 days ago
Type
Full time
Salary
Salary not provided

Related skills

sre site reliability grafana prometheus python

πŸ“‹ Description

  • Define SLIs/SLOs for ~70 components with ownership and error budgets.
  • Build telemetry and reliability platform for cloud and agent-based services.
  • Instrument Python, Go, and Rust components for reliability.
  • Consolidate dashboards; retire non-value tooling.
  • Establish alerting with multi-window burn-rate; ensure actionable runbooks.
  • Lead blameless postmortems and on-call improvements across time zones.

🎯 Requirements

  • Production engineering or SRE with SLO framework experience.
  • Strong Python; reading/modifying Go or Rust for instrumentation.
  • Hands-on time-series telemetry: Prometheus/OpenMetrics, Grafana, Alertmanager, ClickHouse or
  • Experience debugging distributed systems beyond Kubernetes.
  • CI/CD tooling experience: Ansible, GitLab CI, Jenkins, etc.
  • Telemetry for non-scrapable systems; privacy considerations on customer infra.

🎁 Benefits

  • Fully remote work with global flexibility.
  • 25 days vacation, paid holidays, unlimited sick leave.
  • Private medical insurance contributions; coworking space reimbursement.
  • Gym/sports reimbursement; professional development opportunities.
  • Opportunity to define SRE function and operational culture.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’