Software Engineer, Data Infrastructure

Added
8 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

s3 python kubernetes go csi drivers

📋 Description

  • Design, build, and operate the distributed storage system that feeds model training and evaluation.
  • Run this system multiple on Kubernetes clusters at petabyte scale.
  • Work with researchers and training-infra teams on how jobs actually read and write data, and turn
  • Work through the networking, I/O, and consistency problems of moving large datasets and checkpoints

🎯 Requirements

  • Strong storage fundamentals, including replication, consistency, caching, and data lifecycle
  • Strong coding ability. We work in Python and Go; experience in either is enough, but you should be
  • Experience running stateful systems on Kubernetes, including Persistent Volumes, CSI drivers, and
  • Hands-on experience with cloud object storage such as S3 as well as POSIX-style filesystems.
  • Experience with parallel or HPC filesystems such as Weka, VAST, or Lustre (bonus).
  • Familiarity with the data-loading and checkpointing patterns used in large-scale model training

🎁 Benefits

  • A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch.
  • Full health and dental benefits, including a separate budget for mental health.
  • RRSP matching, 401K, Pension Scheme.
  • 100% Parental Leave top-up for up to 6 months, for either parent.
  • Annual enrichment benefits: Arts & culture, fitness/wellness, quality time, and a workspace
  • 6 weeks of paid vacation (30 working days!).
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →