Added
16 days ago
Type
Full time
Salary
Salary not provided

Related skills

datadog docker terraform aws kubernetes

πŸ“‹ Description

  • Define and implement reliability strategy across the platform (SLOs/SLIs, incident practices
  • Drive architectural decisions to keep systems scalable, resilient, and observable as AI workloads
  • Design and own event-driven messaging, transitioning from synchronous to durable asynchronous
  • Manage AWS infrastructure with IaC for provisioning, configuration, deployment, and operations.
  • Ensure Kubernetes and containerized workloads scale reliably with increasing load.
  • Build and maintain observability with monitoring, dashboards, APM, and distributed tracing.

🎯 Requirements

  • Extensive SRE/Platform Engineering/DevOps experience with production-scale systems ownership.
  • Deep expertise in event-driven architectures and messaging (Kafka, NATS, RabbitMQ) and related
  • Strong AWS expertise (EC2, VPC, IAM, S3, RDS) plus solid networking fundamentals.
  • Hands-on IaC experience with Terraform, Pulumi, or similar tools with Git workflows.
  • Strong Kubernetes and Docker production experience (lifecycle, resource limits, health checks
  • Observability expertise with Datadog or equivalent (dashboards, monitoring, APM, tracing, alerting).

🎁 Benefits

  • Competitive compensation.
  • Fully remote with flexible locations.
  • One-time home office allowance.
  • Company-provided equipment.
  • Stock options.
  • Health plan available worldwide.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’