Senior Site Reliability Engineer

Added
22 days ago
Type
Full time
Salary
Salary not provided

Related skills

docker ansible terraform sql python

πŸ“‹ Description

  • Collaborate to design scalable, secure, highly available systems.
  • Establish and manage SLOs and SLAs for ClickHouse Cloud.
  • Ensure monitoring and alerting for timely incident detection.
  • Improve incident response and post-mortem analyses.
  • Continuously improve reliability and performance of services.
  • Plan, enable, and drive Chaos initiatives across teams.
  • Manage on-call processes and escalation best practices.

🎯 Requirements

  • Bachelor's or Master's in Computer Science or related field.
  • 8+ years in Site Reliability Engineering or related field.
  • Production experience with ClickHouse.
  • Hands-on with Go and/or Python.
  • Strong knowledge of AWS, Azure, or Google Cloud Platform.
  • Expertise in distributed databases and SQL.
  • Experience with Kubernetes or Docker Swarm.
  • Automation using Ansible, Terraform, or Puppet.

🎁 Benefits

  • Flexible, remote-friendly environment; works in 20+ countries.
  • Healthcare coverage with employer contributions.
  • Equity via stock options.
  • Flexible time off (US) / generous elsewhere.
  • $500 home office setup for remote staff.
  • Global company gatherings and offsites.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’