Senior Site Reliability Engineer (SRE, Compute Node Team)

Added
7 hours ago
Type
Full time
Salary
Salary not provided

Related skills

linux kubernetes virtualization containers ebpf

๐Ÿ“‹ Description

  • Ensure reliability, availability, and performance of compute nodes across clouds.
  • Analyze and troubleshoot complex Linux systems (user/kernel space).
  • Investigate and resolve production issues with CPU, memory, NUMA, cgroups, and scheduling.
  • Hands-on virtualization work with QEMU/KVM and Linux-native solutions.
  • Design observability: metrics, logs, traces, alerts, SLIs, SLOs.
  • Lead incident response, perform root-cause analysis, and drive post-incident improvements.

๐ŸŽฏ Requirements

  • Deep Linux expertise (user/kernel space, scheduling, memory, filesystems, cgroups, namespaces).
  • Strong understanding of system architecture and performance trade-offs across infra layers.
  • Hands-on virtualization experience (QEMU/KVM), VM lifecycle, and performance optimization.
  • Experience with containers, namespaces, and resource isolation.
  • Strong debugging and hypothesis-driven incident investigation skills.
  • Experience building observability solutions.

๐ŸŽ Benefits

  • Competitive compensation package.
  • Flexible environment with project ownership.
  • Career growth and continuous learning.
  • Work on AI infrastructure projects.
  • Collaborative international engineering teams.
  • Solve challenging problems at scale.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’