Lead Site Reliability Engineer - Imunify Reliability Platform

Added
5 days ago
Type
Full time
Salary
Salary not provided

Related skills

rust grafana prometheus python go

πŸ“‹ Description

  • Define SLIs/SLOs and ownership across ~70 components.
  • Build telemetry for customer-hosted agents and cloud services.
  • Instrument Python, Go, and Rust components.
  • Consolidate dashboards and enable reliable observability tooling.
  • Establish incident response and blameless postmortems.

🎯 Requirements

  • Extensive production SRE/engineering experience with SLO framework.
  • Strong Python; ability to modify Go or Rust for instrumentation.
  • Hands-on telemetry at scale: Prometheus/OpenMetrics, Grafana, Alertmanager, ClickHouse or similar.
  • Experience debugging distributed systems beyond Kubernetes-centric view.
  • Practical config mgmt and CI/CD tooling (Ansible, GitLab CI, Jenkins, etc).
  • Telemetry for unscannable systems: push-based collection, privacy considerations.

🎁 Benefits

  • Fully remote work with flexible hours worldwide.
  • 25 days vacation, 10 holidays, unlimited sick leave.
  • Private medical insurance contribution; coworking space reimbursement.
  • Gym/sports reimbursement; professional development opportunities.
  • Opportunity to define SRE function and reliability culture.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’