Added
28 days ago
Type
Full time
Salary
Salary not provided

Related skills

datadog terraform aws kubernetes rabbitmq

πŸ“‹ Description

  • Define reliability strategy with SLOs/SLIs and incident practices
  • Drive scalable, observable architectures and incident response
  • Design event-driven messaging and asynchronous patterns
  • Manage AWS infra with IaC and Kubernetes at scale
  • Build comprehensive observability with monitoring and tracing
  • Lead major incidents and drive permanent improvements

🎯 Requirements

  • Extensive SRE/Platform/DevOps with production-scale systems
  • Deep expertise in event-driven architectures using Kafka/NATS/RabbitMQ
  • Strong AWS (EC2/VPC/IAM/S3/RDS) and networking
  • Terraform/Pulumi IaC with version-controlled workflows
  • Kubernetes/Docker production experience with observability
  • Experience defining SLOs/SLIs and chaos engineering

🎁 Benefits

  • Fully remote, home office allowance
  • Company-provided equipment
  • Stock options
  • Health plan
  • Flexible days off
  • Opportunity to work on globally scaled infra
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’