Lead Site Reliability Engineer - Imunify Reliability Platform

Added
5 days ago
Type
Full time
Salary
Salary not provided

Related skills

rust ansible grafana prometheus python

πŸ“‹ Description

  • Define SLIs/SLOs for ~70 components with ownership, measurement, and error budgets.
  • Build a reliability taxonomy covering availability, latency, telemetry health, and security
  • Design telemetry pipelines for customer-hosted agents and cloud services.
  • Extend instrumentation across Python, Go, and Rust components.
  • Consolidate dashboards and establish a reliable observability platform.
  • Implement SLO-driven, symptom-based alerting with clear runbooks.

🎯 Requirements

  • Production engineering or SRE experience with SLO framework definition.
  • Strong Python; ability to read/modify Go or Rust for instrumentation.
  • Hands-on time-series/event telemetry exp. with Prometheus/OpenMetrics, Grafana, Alertmanager
  • Experience debugging distributed systems beyond Kubernetes-centric ops.
  • Practical config mgmt/CI-CD tooling (Ansible, GitLab CI, Jenkins).
  • Understand telemetry for systems with push-based collection, privacy concerns on customer infra.

🎁 Benefits

  • Fully remote with worldwide flexibility.
  • 24 vacation days; 10 holidays; unlimited sick leave.
  • Private medical insurance contribution; co-working space and gym reimbursements.
  • Professional development and opportunities to patent ideas.
  • Remote-first, asynchronous work across time zones.
  • Opportunity to define SRE function and operational culture from the ground up.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’