Lead Site Reliability Engineer - Imunify Reliability Platform

Added
20 hours ago
Type
Full time
Salary
Salary not provided

Related skills

rust grafana prometheus python go

πŸ“‹ Description

  • Define SLIs for ~70 components with ownership, measurement, SLOs, and error budgets.
  • Develop a reliability taxonomy for availability, latency, telemetry, and security control efficacy.
  • Ensure reliability indicators are independent and not bypassable by failures they detect.
  • Design telemetry collection for customer-hosted agents and cloud services; balance push-based
  • Extend instrumentation across Python, Go, and Rust components with product teams.
  • Consolidate dashboards into a reliable observability platform; retire non-value tooling.

🎯 Requirements

  • Substantial production engineering or SRE experience with SLO framework ownership.
  • Strong Python skills; ability to modify Go or Rust for instrumentation.
  • Experience with time-series/event telemetry at scale (Prometheus/OpenMetrics, Grafana
  • Hands-on debugging of distributed systems beyond Kubernetes-centric ops.
  • Experience with production CI/CD tooling (Ansible, GitLab CI, Jenkins).
  • Understanding of telemetry for systems with push-based collection, sampling, and privacy

🎁 Benefits

  • Fully remote work with flexible hours (worldwide).
  • 24 vacation days/year; 10 national holidays; unlimited sick leave.
  • Private medical insurance contribution; coworking space reimbursement; gym rebate.
  • Professional development through challenging projects and mentoring.
  • Opportunity to earn a patent-worthy idea reward.
  • Remote-first, asynchronous environment across time zones.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’