Production Engineer - Team Lead

Added
14 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

terraform bash aws grafana prometheus

πŸ“‹ Description

  • Act as Incident Commander during incidents and guide rapid resolution.
  • Coordinate cross-functional teams during incidents.
  • Lead root cause analysis (RCA) and implement durable fixes.
  • Own post-incident reviews and actionable learnings.
  • Develop incident response playbooks and escalation processes.
  • Define and track SLOs aligned with business goals.

🎯 Requirements

  • 4+ years in production engineering, cloud ops, SRE, or incident response.
  • Kubernetes-based infrastructure, AWS, and GCP.
  • ITIL and SRE best practices.
  • Prometheus and Grafana monitoring; strong observability.
  • Automation with Python, Bash, Terraform.
  • Decision-making under pressure; strong communication.
  • Experience mentoring and coaching technical teams.

🎁 Benefits

  • Medical, dental, and vision insurance.
  • 401(k) with employer match.
  • Flexible PTO and parental leave.
  • Health Savings Account and FSA.
  • Tuition reimbursement and ESPP.
  • Mental wellness benefits.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’