Site Reliability Engineer 3

Added
2 hours ago
Type
Full time
Salary
Salary not provided

Related skills

azure ansible terraform linux aws

๐Ÿ“‹ Description

  • End-to-end reliability ownership for production systems (on-call, incidents, SLO/SLI).
  • Build and maintain observability across metrics, logs, traces with anomaly detection.
  • Design AIOps workflows for event ingestion, enrichment, correlation, deduplication, noise reduction.
  • Develop AI-assisted operational processes (log analysis, incident summaries, RCA, runbooks).
  • Create automation solutions with AI-enhanced runbooks and self-healing workflows.
  • Implement safeguards, approvals, rollbacks, and audit controls for automation.

๐ŸŽฏ Requirements

  • 6+ years in SRE, AIOps, DevOps, or production engineering in large-scale cloud environments.
  • Strong knowledge of Linux/Unix, networking, distributed systems, and AWS/Azure/Google Cloud.
  • Expert-level experience with ELK/OpenSearch: log ingestion, indexing, dashboards, and alerting.
  • Strong observability across logs, metrics, and distributed tracing.
  • Experience preparing telemetry data for AIOps: enrichment, service mapping, normalization, and correlation.
  • Hands-on AIOps: anomaly detection, alert correlation, intelligent alerting, and automated enrichment.

๐ŸŽ Benefits

  • Competitive compensation package.
  • Fully remote work opportunity from India.
  • Opportunity to work with advanced cloud, automation, and AI-driven technologies.
  • Exposure to large-scale production environments and complex reliability challenges.
  • Collaboration with globally distributed engineering teams.
  • Professional growth opportunities through innovative technology initiatives.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’