Added
20 hours ago
Type
Full time
Salary
Salary not provided

Related skills

terraform grafana prometheus python application insights

πŸ“‹ Description

  • Own tools, processes, and systems for visibility into AI applications, platform services, and
  • Design, implement, and operate observability solutions for AI workloads, including LLM and agent
  • Configure and maintain AI tracing systems to capture latency, token usage, cost, quality metrics
  • Develop internal tooling and automation solutions using Python for instrumentation and data
  • Build and maintain dashboards, metrics, and monitoring solutions using Grafana, Prometheus, and
  • Instrument AI platforms to provide visibility into health, usage, performance, cost, and SLIs/SLOs.

🎯 Requirements

  • 5–8 years of experience in observability, SRE, platform engineering, DevOps, or cloud engineering
  • Strong hands-on experience with Azure Monitor, Application Insights, Log Analytics, and Managed
  • Experience with Langfuse, Grafana, and Prometheus for AI and application monitoring.
  • Solid knowledge of Terraform and CI/CD practices.
  • Strong Python skills for automation, instrumentation, exporters, and internal tooling development.
  • Familiarity with machine learning workloads and AI-specific observability requirements.

🎁 Benefits

  • Competitive compensation package.
  • Career development and continuous learning opportunities.
  • Flexible working environment with strong ownership and autonomy.
  • Opportunity to work on impactful AI infrastructure and technology projects.
  • Collaborative culture with highly skilled international teams.
  • Chance to contribute to the evolution of next-generation AI platforms.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’