Staff Software Engineer, Site Reliability Engineer

Added
14 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

datadog terraform cloudformation pagerduty python

πŸ“‹ Description

  • Design, implement, and manage monitoring and infrastructure resources across 50+ regions
  • Lead incident management, including postmortems and root cause analyses
  • Automate tasks and workflows for capacity planning and safe rollouts
  • Establish security, compliance, and reliability practices across the software lifecycle
  • Optimize infrastructure costs via capacity planning and build-vs-buy decisions
  • Provide technical mentorship and foster team growth

🎯 Requirements

  • 10+ years in SRE or similar roles supporting production systems
  • Experience with IaC tools (Pulumi, Terraform, CloudFormation, etc.)
  • Familiar with observability tools (Datadog, Sentry) and incident response (PagerDuty, IncidentIO)
  • Proficient with cloud platforms (Azure, GCP, AWS)
  • Strong programming skills (Python, Bash, Go)
  • Proven ability to diagnose complex problems and implement durable solutions
  • Solid CI/CD, Kubernetes, containerization, networking, databases, and cloud security
  • Excellent problem-solving and attention to detail

🎁 Benefits

  • In-person role in San Francisco with relocation assistance

🚚 Relocation support

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’