Added
7 days ago
Type
Full time
Salary
Salary not provided

Related skills

terraform grafana prometheus kubernetes opentelemetry

๐Ÿ“‹ Description

  • Operate and improve the shared telemetry platform for metrics, logs, traces, alerting, dashboards
  • Maintain metrics collection, storage, querying, dashboards, and alerting with Prometheus stack
  • Manage log pipelines using Vector, Splunk, and Loki; ensure reliability and throughput.
  • Operate tracing and profiling with Grafana Alloy, Tempo, OpenTelemetry, Pyroscope.
  • Deploy telemetry services with Terraform, Terragrunt, and multi-env orchestration.
  • Troubleshoot missing data, slow queries, broken alerts, capacity issues.

๐ŸŽฏ Requirements

  • 3+ years in SRE/Platform/Observability or similar production role.
  • Experience with large-scale telemetry data (metrics/logs/traces).
  • Prometheus or Prometheus-compatible stack experience.
  • Distributed systems troubleshooting; data flow and capacity issues.
  • Infra as Code experience, esp. Terraform, and CI/CD.
  • Experience with container workloads (Nomad, Kubernetes or similar).

๐ŸŽ Benefits

  • Culture-driven, globally minded tech environment.
  • On-call and incident learnings shape platform improvements.
  • Commitment to diversity and equal opportunity hiring.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’