Senior Observability & Telemetry Engineer - Radian Arc

Added
23 hours ago
Type
Full time
Salary
Salary not provided

Related skills

rust grafana prometheus python go

πŸ“‹ Description

  • Design, implement, and operate scalable telemetry pipelines for metrics, logs, and traces across
  • Architect telemetry storage for time-series and event data; define instrumentation and SLIs/SLOs.
  • Build observability across compute, storage, networking, GPU, and AI workloads; identify
  • Develop dashboards and monitoring tools for workload health and performance insights.
  • Build network/infra telemetry using Python or Go; integrate gNMI, SNMP, and streaming telemetry.
  • Develop alerting and anomaly detection; integrate signals into incident workflows.

🎯 Requirements

  • Proven experience operating observability systems at production scale.
  • Strong Go, Python, or Rust programming skills.
  • Hands-on with Prometheus, OpenTelemetry, Grafana; large telemetry DBs like ClickHouse.
  • Experience with GPU cloud/HPC/AI infra and distributed training/inference monitoring.
  • Knowledge of NVIDIA DCGM, NVML, GPU telemetry and AI workload metrics.
  • Strong networking/infra telemetry experience (gNMI, SNMP, streaming telemetry).

🎁 Benefits

  • Attractive compensation; permanent, full-time; EMEA-based remote.
  • Flexible/hybrid-friendly environment for international collaboration.
  • Work on GPU/AI/cloud infra challenges; exposure to advanced observability tech.
  • Opportunity to influence platform standards and mentor engineers.
  • Career growth in a fast-growing international scale-up.
  • Inclusive, diverse environment with fair development support.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’