Lead Site Reliability Engineer - Imunify Reliability Platform

Added
6 days ago
Type
Full time
Salary
Salary not provided

Related skills

rust ansible grafana prometheus python

πŸ“‹ Description

  • Define SLIs/SLOs and ownership for ~70 components across cloud services and agents.
  • Build telemetry and reliability platform; improve detection of silent degradations.
  • Extend instrumentation to Python, Go, and Rust codebases.
  • Create a scalable alerting system with clear runbooks and ownership.
  • Collaborate with engineering leads to shape reliability culture and foundations.

🎯 Requirements

  • Substantial production engineering or SRE experience with SLO framework implementation.
  • Strong Python skills; ability to modify Go or Rust for instrumentation.
  • Experience with time-series telemetry at scale (Prometheus/OpenMetrics, Grafana, Alertmanager).
  • Hands-on with distributed systems on bare metal/long-lived hosts, not just Kubernetes.
  • Experience with CI/CD tools (Ansible, GitLab CI, Jenkins).
  • Knowledge of telemetry for non-scrapable systems, privacy, and partial reporting.

🎁 Benefits

  • Fully remote work with flexible hours worldwide.
  • 24 vacation days; 10 paid holidays; unlimited sick leave.
  • Private medical insurance contribution, coworking space reimbursement.
  • Gym/sports reimbursement; professional development opportunities.
  • Opportunity to define an SRE function and reliability culture from the ground up.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’