Added
17 days ago
Type
Full time
Salary
Salary not provided

Related skills

datadog terraform aws python kubernetes

πŸ“‹ Description

  • Define and implement the reliability strategy across the platform, including SLOs, SLIs, error
  • Drive major architectural decisions as infrastructure evolves, evaluating technologies and
  • Design and own event-driven communication and messaging infrastructure, including the transition
  • Manage and evolve cloud infrastructure on AWS, using Infrastructure as Code to automate
  • Ensure Kubernetes and containerized workloads scale reliably as transaction volumes and AI
  • Build and maintain comprehensive observability through monitoring, dashboards, alerting

🎯 Requirements

  • Extensive experience in Site Reliability Engineering, Platform Engineering, DevOps, or a closely
  • Deep expertise in event-driven architecture and messaging systems such as Kafka, NATS, or RabbitMQ
  • Strong AWS expertise across services such as EC2, VPC, IAM, S3, and RDS, combined with solid
  • Hands-on Infrastructure as Code experience using Terraform, Pulumi, or similar tools, with
  • Strong production experience with Kubernetes and Docker, including container lifecycle management
  • Proven observability expertise using Datadog or equivalent platforms, including dashboards

🎁 Benefits

  • Competitive compensation.
  • Fully remote working environment with the flexibility to work from different locations.
  • One-time home office allowance to help create an effective workspace.
  • Company-provided work equipment.
  • Stock options.
  • Health plan available wherever you are.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’