Site Reliability Engineer

Added
3 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

aws grafana prometheus kubernetes typescript

πŸ“‹ Description

  • Design and evolve reliability and performance systems for WorkOS
  • Collaborate with product/infra to ensure prod-ready, observable services
  • Define and measure SLIs/SLOs to guide reliability
  • Write and optimize backend systems in TypeScript
  • Improve incident response and drive postmortems
  • Develop internal tools to operate and scale our systems

🎯 Requirements

  • Experience operating and scaling production systems in AWS
  • Familiarity with monitoring, alerting, incident response, and RCA
  • Comfort across compute, networking, storage, observability tooling
  • Strong debugging and systems thinking across services
  • Ability to work independently, take ownership, and drive projects
  • Nice to have: Kubernetes, OpenTelemetry, observability stacks

🎁 Benefits

  • 401k matching
  • Competitive equity
  • Healthcare, dental, and vision coverage
  • Parental leave and paid time off
  • Wellness and fitness stipends
  • Commuter benefits for SF/NYC
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’