Added
20 hours ago
Type
Full time
Salary
Salary not provided

Related skills

terraform grafana prometheus python kubernetes

πŸ“‹ Description

  • Own tools, processes, and systems for AI observability across applications, platforms, and infra.
  • Design, implement, and operate observability solutions for AI workloads (LLMs and agents).
  • Configure AI tracing to capture latency, token usage, cost, quality, prompt analytics, model
  • Develop internal tooling using Python for instrumentation and data collection.
  • Build dashboards and monitoring with Grafana, Prometheus, and cloud observability platforms.
  • Instrument AI platforms to provide health, usage, performance, and SLO visibility.

🎯 Requirements

  • 5–8 years in observability, SRE, DevOps, or cloud engineering roles.
  • Hands-on experience with Azure Monitor, Application Insights, Log Analytics, and Grafana.
  • Experience with Langfuse, Grafana, and Prometheus for AI/app monitoring.
  • Strong Terraform and CI/CD skills.
  • Strong Python for automation, instrumentation, exporters, and tooling.
  • Familiarity with ML workloads and AI observability requirements.

🎁 Benefits

  • Competitive compensation package.
  • Career development and continuous learning opportunities.
  • Flexible working environment with ownership and autonomy.
  • Opportunity to work on impactful AI infrastructure projects.
  • Collaborative culture with international teams.
  • Chance to contribute to next-gen AI platforms.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’