Senior Site Reliability Engineer

Added
15 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

terraform cloudformation aws prometheus python

πŸ“‹ Description

  • Build monitoring to ensure platform health and reliability.
  • Create alerts and runbooks for faster issue detection and remediation.
  • Debug complex cross-component issues and implement fixes to prevent recurrence.
  • Participate in on-call rotations and blameless postmortems.
  • Design and implement platform components to improve customer workflows.
  • Build Kubernetes controllers to automate operations.

🎯 Requirements

  • BS in CS/EE/Robotics or related field +4 yrs exp (MS 2+; PhD optional).
  • Strong Linux internals, TCP/IP networking, storage subsystems.
  • Hands-on Go or Python development for production systems.
  • Cloud experience with AWS and/or GCP; IaC via Terraform/CloudFormation.
  • Kubernetes: controllers in Go; production-grade Kubernetes.
  • Metrics/monitoring with Prometheus; tracing with Jaeger/Tempo.
  • Focus on reliability via SLOs and scalable designs for budgets.
  • Strong communication in diverse, distributed teams.

🎁 Benefits

  • Competitive compensation.
  • Comprehensive health, dental, and vision insurance.
  • HSA with employer match.
  • Employer-matched 401(k) with immediate vesting.
  • Paid parental and medical leave; unlimited vacation and 15 holidays.
  • Daily lunches, snacks, and beverages at all offices.

πŸ›ƒ Visa sponsorship

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’