Added
16 days ago
Type
Full time
Salary
Salary not provided

Related skills

datadog jira confluence splunk nagios

๐Ÿ“‹ Description

  • Own key parts of incident, event, and problem management.
  • Lead or co-lead restoration for outages with cross-functional teams.
  • Provide 24/7 coverage via on-call rotation for live incidents.
  • Communicate updates to stakeholders during incidents and post-incident reviews.
  • Collaborate with Engineering, Infrastructure, Product, and Support.
  • Aim for high uptime in cloud, hybrid, and on-prem environments.

๐ŸŽฏ Requirements

  • 1-3 years Incident Commander level responsibility.
  • Ability to craft concise exec-ready outage docs.
  • Strong monitoring/observability knowledge for health assessment.
  • Interest in AIOps: AI-driven alerting, anomaly detection, auto-remediation.
  • Comfort using AI tools (Claude, ChatGPT, Copilot) in workflows.
  • Experience with incident/problem management processes and tools (Jira, Confluence, Jira Service

๐ŸŽ Benefits

  • Medical insurance coverage for employees and dependents.
  • Group Term & Group Personal Accident Insurance.
  • Work-Life Balance: generous leave policy and holidays.
  • Financial Security: Provident Fund & Gratuity.
  • Wellness programs and ongoing learning opportunities.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’