Lead Site Reliability Engineer - Imunify Reliability Platform

Added
5 days ago
Type
Full time
Salary
Salary not provided

Related skills

rust grafana prometheus python go

πŸ“‹ Description

  • Define SLIs, SLOs for ~70 components with ownership and error budgets.
  • Develop reliability taxonomy for availability, latency, telemetry health.
  • Ensure indicators are measurable and not spoofable by failures.
  • Build telemetry pipeline for customer-hosted agents and cloud services.
  • Extend instrumentation across Python, Go, Rust components.
  • Consolidate dashboards and reporting into a lean observability platform.

🎯 Requirements

  • Production engineering or SRE experience with SLO framework setup.
  • Strong Python; able to read/modify Go or Rust code for instrumentation.
  • Hands-on time-series/event telemetry at scale: Prometheus/OpenMetrics, Grafana, Alertmanager
  • Experience debugging distributed systems beyond Kubernetes-centric ops.
  • Practical config mgmt and CI/CD tooling (Ansible, GitLab CI, Jenkins).
  • Telemetry for systems with push-based collection, privacy, and partial reporting.

🎁 Benefits

  • Fully remote work with global flexibility.
  • Professional development through challenging projects and mentoring.
  • Remote-first, asynchronous environment across time zones.
  • Opportunity to define an SRE function and reliability culture.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’