Site Reliability Engineer (SRE)

Added
17 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

sre docker scripting grafana prometheus

๐Ÿ“‹ Description

  • Monitor platform health across multiple client environments using tools like Grafana and Prometheus
  • Respond to and triage incidents, following established runbooks and escalation paths
  • Participate in post-incident reviews and contribute to postmortem documentation
  • Support the Service Desk team with technical triaging, incident classification, and resolution
  • Maintain and improve observability dashboards, alerts, and SLI, and SLO tracking
  • Write and maintain runbooks, operational documentation, and knowledge base articles

๐ŸŽฏ Requirements

  • 2-3 years of experience in SRE, platform operations, or a technical Service Desk role
  • Strong proficiency in English (written and verbal)
  • Experience with monitoring and observability tools such as Grafana, Prometheus, or equivalent
  • Solid understanding of incident management processes (triaging, escalation, postmortems)
  • Experience supporting e-commerce platforms
  • Scripting skills in shell and/or Python for automation and operational tasks

๐ŸŽ Benefits

  • Generous vacation policy
  • Flexible work arrangements
  • Training budgets
  • Mentorship
  • AI and strategic upskilling
  • Inclusive and safe environment
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’