Lead Site Reliability Engineer - Imunify Reliability Platform

Added
5 days ago
Type
Full time
Salary
Salary not provided

Related skills

rust grafana prometheus python kubernetes

πŸ“‹ Description

  • Lead SRE for Imunify Reliability Platform across 70 components.
  • Define SLIs/SLOs, error budgets, and ownership across squads.
  • Build telemetry and reliability platform for cloud services and agents.
  • Extend instrumentation in Python, Go, and Rust.
  • Consolidate dashboards and implement alerting with clear ownership and runbooks.
  • Shape reliability culture and operational standards across time zones.

🎯 Requirements

  • Production engineering or SRE experience with SLO framework definition.
  • Strong Python; reading/modifying Go or Rust for instrumentation.
  • Hands-on time-series/event telemetry at scale (Prometheus/OpenMetrics, Grafana, Alertmanager).
  • Experience debugging distributed systems beyond Kubernetes-centric ops.
  • Experience with CI/CD tools (Ansible, GitLab CI, Jenkins).
  • Knowledge of telemetry under restricted/privacy-conscious environments.

🎁 Benefits

  • Fully remote work with flexible hours.
  • 24 vacation days, 10 holidays, unlimited sick leave.
  • Private medical insurance support, coworking space reimbursement, gym reimbursement.
  • Professional development and opportunities to influence SRE culture.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’