Senior Site Reliability Engineer (SRE, Compute Node Team)

Added
5 hours ago
Type
Full time
Salary
Salary not provided

Related skills

linux kubernetes virtualization observability qemu/kvm

๐Ÿ“‹ Description

  • Ensure reliability, availability, and performance of compute nodes running VMs.
  • Analyze and troubleshoot Linux systems across user and kernel space.
  • Investigate and resolve CPU, memory, NUMA, cgroups, and scheduling issues.
  • Hands-on with virtualization tech, including QEMU/KVM and Linux virtualization.
  • Design observability at infrastructure level: metrics, logs, traces, alerts, SLIs, SLOs.
  • Lead incident response, perform root-cause analysis, drive improvements.

๐ŸŽฏ Requirements

  • Deep Linux expertise: user/kernel space, cgroups, namespaces.
  • Strong system architecture and performance trade-off understanding.
  • Hands-on virtualization with QEMU/KVM; VM lifecycle and perf.
  • Container tech: namespaces and resource isolation.
  • Debugging with hypothesis-driven incident investigations.
  • SRE principles: reliability, system design, ownership.

๐ŸŽ Benefits

  • Competitive compensation package.
  • Flexible working environment with ownership over projects.
  • Career growth and continuous learning.
  • Work on impactful AI infrastructure projects.
  • Collaborative international engineering teams.
  • Solve challenging technical problems at scale.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’