Senior Site Reliability Engineer - Storage

Added
8 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

datadog linux grafana prometheus ceph
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now β†’

πŸ“‹ Description

  • Join Lambda's Storage Engineering team.
  • Maintain reliability and performance of production storage.
  • Work across data centers with software-defined storage.
  • Collaborate on automation, CI/CD, and incident response.

🎯 Requirements

  • 5+ years Linux production/HPC experience.
  • Hands-on SDS at scale; CEPH/Lustre/GPFS or similar.
  • Monitoring/logging: Prometheus, Grafana, Alertmanager, Datadog, SumoLogic.
  • Kubernetes with GitOps (ArgoCD, Helm/Kustomize).
  • CI/CD tooling (GitHub Actions, Jenkins, BuildKite); Python or Go.
  • Terraform/Ansible; NIC and storage protocol knowledge.

🎁 Benefits

  • Competitive salary and equity.
  • Health/dental/vision coverage.
  • Wellness and commuter stipends.
  • 401k with 2% match (USA).
  • Flexible PTO.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’