Site Reliability Engineer - Data Platform

Added
2 days ago
Type
Full time
Salary
Salary not provided

Related skills

site reliability java ansible puppet linux

📋 Description

  • Design, implement and operate our data platforms.
  • Improve observability so we catch issues before our users do.
  • Build automation to reduce toil and allow our systems to scale.
  • Support and own reliability of critical services (e.g. HDFS, Kafka and Dremio).
  • Drive long-term architectural improvements, not just fixing issues, but preventing them.

🎯 Requirements

  • Strong experience managing distributed data platforms (e.g. Kafka, Hadoop, Spark, Dremio)
  • Hands on experience deploying, configuring and orchestrating software on Linux and Kubernetes, with
  • Strong experience with infrastructure as code (Ansible preferred) and best practices.
  • Proficient programming experience in Python.
  • Ability to read, write and tune SQL queries.
  • Comfortable reading Java source code, tuning and debugging running JVMs.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →