Added
20 hours ago
Type
Full time
Salary
Salary not provided

Related skills

terraform grafana prometheus python ci/cd

πŸ“‹ Description

  • Own tools and systems for visibility into AI apps, platforms, and infra performance
  • Design, implement, operate observability for AI workloads including LLMs
  • Configure AI tracing systems for latency, token usage, cost, quality, prompts, model versions
  • Build automation tooling in Python for instrumentation and data collection
  • Maintain dashboards and monitoring with Grafana, Prometheus, cloud observability
  • Instrument AI platforms for health, usage, cost, and SLOs

🎯 Requirements

  • 5–8 years in observability, SRE, DevOps, or cloud engineering
  • Hands-on with Azure Monitor, Application Insights, Log Analytics, Grafana
  • Experience with Langfuse, Grafana, Prometheus for AI monitoring
  • Terraform and CI/CD practices
  • Strong Python for automation and tooling
  • Familiarity with ML workloads and AI observability needs

🎁 Benefits

  • Competitive compensation
  • Career development and learning opportunities
  • Flexible working environment with ownership
  • Opportunity to work on impactful AI infrastructure
  • Collaborative international teams
  • Contribute to next-gen AI platforms
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’