Technical Lead, Platform Engineering (Observability)

Added
14 days ago
Type
Full time
Salary
Salary not provided

Related skills

datadog grafana prometheus python kubernetes

๐Ÿ“‹ Description

  • Lead observability platform design, build, and operation for metrics, logs, traces, dashboards, and
  • Define engineering standards, drive architectural decisions, and oversee design-to-operations
  • Improve MTTD/MTTM and overall service reliability and resilience.
  • Develop AI-assisted anomaly detection, alert correlation, and root cause analysis capabilities.
  • Build self-service tooling for developers to instrument services, create dashboards, and configure
  • Define SLOs/SLIs, error budgets, and instrumentation practices across teams.

๐ŸŽฏ Requirements

  • 9+ years in platform engineering, SRE, or related fields with scalable production systems.
  • Hands-on observability with Datadog, Prometheus, Grafana, or equivalents.
  • Deep understanding of metrics, logs, tracing, and observability data pipelines.
  • Proven Kubernetes experience in production environments.
  • Strong Go or Python for tooling and automation.
  • Experience with GCP and/or AWS, Terraform, and IaC.

๐ŸŽ Benefits

  • Full-time employment with a hybrid/remote-friendly setup, Bengaluru-based collaboration optional.
  • Opportunity to shape observability practices across a large distributed system.
  • Exposure to cloud infra, Kubernetes, AI-powered operations, and large-scale distributed systems.
  • Strong focus on technical leadership, mentorship, and career growth.
  • Collaborative environment with platform, SRE, security, and product teams.
  • Potential to directly improve productivity, reliability, and incident response at scale.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’