Added
1 hour ago
Type
Full time
Salary
Salary not provided

Related skills

prometheus elasticsearch capacity planning observability engineering
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now β†’

πŸ“‹ Description

  • Lead, hire, onboard, and develop a distributed engineering team working asynchronously.
  • Set priorities with Site Reliability Engineering, Product Engineering, and GitLab Dedicated teams
  • Own the reliability, scalability, and cost of the team's metrics, logging, alerting, and capacity
  • Reduce noisy or missing alerts and telemetry gaps, and use SLOs, error budgets, and self-service
  • Guide technical decisions about time-series storage, high-cardinality metrics, log pipelines, and
  • Participate in the Incident Manager On Call (IMOC) rotation, coordinating the response to

🎯 Requirements

  • Experience leading an observability, platform engineering, or site reliability engineering team
  • Technical knowledge of metrics systems such as Prometheus and long-term storage, logging platforms
  • Experience using SLOs, error budgets, and capacity forecasts to make reliability and investment
  • Experience operating a large software-as-a-service platform and investigating production issues
  • Experience participating in and improving production on-call rotations, including incident
  • The ability to explain technical tradeoffs to engineering partners and other stakeholders.

🎁 Benefits

  • Benefits to support your health, finances, and well-being
  • Flexible Paid Time Off
  • Team Member Resource Groups
  • Equity Compensation & Employee Stock Purchase Plan
  • Growth and Development Fund
  • Parental Leave
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’