Reliability Monitoring Engineer

Added
1 day ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

datadog linux grafana prometheus python

πŸ“‹ Description

  • Design and maintain observability across metrics, logs, traces, dashboards, and alerts.
  • Build scalable telemetry pipelines and monitoring infrastructure.
  • Configure and optimize Prometheus, Grafana, and other tools.
  • Develop monitoring strategies to improve visibility and reliability.
  • Implement distributed tracing, structured logging, and OpenTelemetry.
  • Collaborate with engineering teams to translate signals into outcomes.

🎯 Requirements

  • Bachelor's degree in Computer Science, Engineering, or related field.
  • 5+ years in SRE, platform engineering, reliability, or observability roles.
  • Hands-on with Prometheus, Grafana, and a major observability platform (Datadog/New Relic).
  • OpenTelemetry, distributed tracing, telemetry pipelines, and structured logging.
  • Proficiency in Go, Python, or Java.
  • Linux, networking, and container technologies.

🎁 Benefits

  • Competitive salary $100k–$150k.
  • Fully remote work within the United States.
  • Full-time employment with growth opportunities.
  • Work on modern reliability, monitoring, and cloud initiatives.
  • Collaborative environment focused on innovation and excellence.
  • Exposure to large-scale systems and advanced observability practices.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’