Staff Site Reliability Engineer

Added
14 days ago
Type
Full time
Salary
Salary not provided

Related skills

azure terraform aws kubernetes ci/cd

πŸ“‹ Description

  • Architect, deploy, operate, and improve scalable, secure production environments (AWS preferred).
  • Lead reliability initiatives across multiple streams and establish SRE practices.
  • Design, migrate, and optimize Kubernetes-based infrastructure, including production hardening and
  • Establish Infrastructure-as-Code standards using Terraform or equivalents.
  • Define and operationalize SLIs, SLOs, error budgets, and reliability practices.
  • Strengthen observability across apps, infra, data pipelines, and ML systems.

🎯 Requirements

  • Extensive hands-on SRE/Production Engineering experience.
  • Proven ability to scale SRE practices in high-growth, distributed environments.
  • Deep expertise with AWS or Azure cloud infra and cloud-native architectures.
  • Strong Kubernetes production experience (migration, scaling, security hardening).
  • Advanced Infrastructure-as-Code with Terraform or equivalent.
  • Experience designing and optimizing end-to-end CI/CD pipelines.

🎁 Benefits

  • Full-time, permanent employment.
  • Fully remote position within European time zones.
  • Opportunity to work on infrastructure for AI and agent-based workloads.
  • Significant technical ownership and influence over reliability practices and architecture.
  • Collaborative environment across cloud, Kubernetes, data platforms, and ML ops.
  • International and distributed working environment with career growth in leadership impact.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’