Senior Site Reliability Engineer - Linux Systems & Application Observability

Added
11 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

cloud linux prometheus distributed systems observability

๐Ÿ“‹ Description

  • Build self-healing, fault-tolerant infrastructure and internal tooling that automates repetitive
  • Run gap analysis across our observability stack to find blind spots in telemetry, logging, and
  • Own scalability work across our HashiCorp Nomad service fabric: capacity planning, load testing
  • Extend our observability stack (Prometheus, Honeycomb, OpenTelemetry) with the instrumentation
  • Set SLOs and error budgets with multi-window burn-rate alerting for critical brokerage flows, once
  • Mentor engineers across teams to build a culture of site reliability champions so the practice

๐ŸŽฏ Requirements

  • Hands-on experience designing fault-tolerant, self-healing distributed systems โ€” not just
  • Deep understanding of one or more: distributed systems, Linux systems, cloud-native architectures
  • Experience running gap analyses on observability/telemetry systems: identifying what's not
  • A track record scaling systems under real production load, including capacity planning and
  • Hands-on experience with OpenTelemetry, Prometheus, and Grafana, with the ability to instrument
  • Strong Linux internals and networking fundamentals, including TCP/IP, UDP/multicast, packet

๐ŸŽ Benefits

  • Performance Bonuses
  • Stock Purchase Options
  • Medical/Vision/Dental Benefits
  • 401k Plan
  • 20 Paid Vacation Days (plus an additional paid vacation day the month of your birthday!)
  • 10 Paid Sick Days
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’