Staff Site Reliability Engineer (AI Platform)

Added
7 days ago
Type
Full time
Salary
Salary not provided

Related skills

ansible terraform github actions helm prometheus

๐Ÿ“‹ Description

  • Maintain and harden AWS infrastructure (EC2, ALB/NLB, WAF, IAM, CloudWatch)
  • Operate and evolve EKS clusters powering Python-based AI services
  • Migrate services to Kubernetes using Terraform and Helm
  • Codify infrastructure with Terraform and automate host-level tasks via Ansible
  • Build and improve CI/CD pipelines with GitHub Actions
  • Own observability: Prometheus, Grafana, alerts, and on-call readiness

๐ŸŽฏ Requirements

  • 5+ years of experience managing Linux in production (Ubuntu, Amazon Linux)
  • Strong experience with Kubernetes (ideally EKS), Helm, and Terraform
  • Comfort with running and debugging Python workloads in containers
  • Solid understanding of networking, IAM, and cloud security best practices
  • Hands-on Nginx experience (Ingress and reverse proxy setups)
  • Excellent communication skills; you can explain complex infra to devs clearly

๐ŸŽ Benefits

  • Hybrid onboarding and relocation support for you and family
  • Comprehensive health insurance for you and family
  • Professional development budget for conferences and online courses
  • Flexible benefits package to tailor perks
  • Hybrid work and generous leave options

๐Ÿšš Relocation support

Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’