Added
28 days ago
Type
Full time
Salary
Salary not provided

Related skills

datadog terraform aws kubernetes rabbitmq

πŸ“‹ Description

  • Define and implement reliability strategy across the platform including SLOs/SLIs
  • Drive architectural decisions for scalable, observable systems
  • Design event-driven messaging and durable async architectures
  • Manage AWS cloud infra with IaC for provisioning & deployment
  • Ensure Kubernetes/container workloads scale with AI workloads
  • Build and maintain comprehensive observability (monitoring, tracing)

🎯 Requirements

  • Extensive SRE/Platform Eng/DevOps exp with production systems
  • Expertise in event-driven arch and messaging systems (Kafka/NATS/RabbitMQ)
  • Strong AWS (EC2, VPC, IAM, S3, RDS) and networking fundamentals
  • Infrastructure as Code with Terraform/Pulumi
  • Kubernetes and Docker prod experience
  • Observability expertise (Datadog, dashboards, tracing)

🎁 Benefits

  • Competitive compensation
  • Fully remote with flexible locations
  • Home office allowance
  • Company-provided equipment
  • Stock options
  • Health plan globally available
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’