Senior Site Reliability Engineer (SRE, Compute Node Team)

Added
1 day ago
Type
Full time
Salary
Salary not provided

Related skills

linux kubernetes kvm observability ebpf

๐Ÿ“‹ Description

  • Ensure reliability, availability, and performance of compute nodes across clouds.
  • Analyze and troubleshoot Linux systems in user and kernel space.
  • Investigate and resolve production issues involving CPU, memory, NUMA, and scheduling.
  • Hands-on with virtualization tech, including QEMU/KVM.
  • Design observability at infra layer: metrics, logs, traces, alerts, SLIs, SLOs.
  • Lead incident response, perform root-cause analysis, and drive improvements.

๐ŸŽฏ Requirements

  • Linux expertise across user/kernel space and subsystems (scheduling, memory, cgroups, namespaces).
  • Strong system architecture knowledge and performance trade-offs across layers.
  • Hands-on with virtualization (QEMU/KVM), VM lifecycle, and perf optimization.
  • Practical experience with containers, namespaces, and resource isolation.
  • Strong debugging with hypothesis-driven incident investigation.
  • Solid grasp of SRE principles: reliability, design, and operational ownership.

๐ŸŽ Benefits

  • Competitive compensation package.
  • Flexible working environment with ownership over projects.
  • Opportunities for career growth and continuous learning.
  • Chance to work on impactful AI infrastructure projects.
  • Collaborative international environment with talented engineering teams.
  • Opportunity to solve challenging technical problems at large scale.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’