Senior Site Reliability Engineer

Added
11 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

java grafana prometheus python ruby

πŸ“‹ Description

  • Define customer-meaningful SLOs and set error budgets for broker flows (order execution and market
  • Develop reliability standards, including SLO methods, error-budget policy, observability guide, and
  • Contribute reliability patterns (circuit breakers, retries, bulkheads, load-shedding) into Ruby
  • Extend observability stack (Prometheus, Honeycomb, OpenTelemetry) as workloads scale on Nomad.
  • Design and run tabletop exercises and fault-injection tests to stress platforms against real-world
  • Mentor engineers to build a lasting site reliability culture across teams.

🎯 Requirements

  • Production-quality coding in Ruby and/or Java, plus Python for automation.
  • Experience embedding SRE practices (SLOs, error budgets, burn-rate alerting).
  • Hands-on OpenTelemetry, Prometheus, Grafana instrumentation.
  • Strong Linux internals and networking (TCP/IP, UDP/multicast, packet capture).
  • On-call production experience and blameless post-incident reviews.
  • Experience influencing standards; HashiCorp Nomad/Consul/Vault is a strong plus.

🎁 Benefits

  • Performance bonuses
  • Stock purchase options
  • Medical/Vision/Dental benefits
  • 401k plan
  • Paid vacation and sick days
  • Gym membership reimbursement
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’