Added
3 hours ago
Type
Full time
Salary
Salary not provided

Related skills

datadog terraform github actions linux python

๐Ÿ“‹ Description

  • Design and improve the platform systems for training, evaluation, and production serving.
  • Build infrastructure and tooling for reliable, scalable ML workloads.
  • Develop internal tools and workflows operable by humans and agents.
  • Shape deployment architecture across research and product environments.
  • Improve scheduling, monitoring, and debugging of GPU and cloud workloads.
  • Drive observability, automation, reliability, and developer experience.

๐ŸŽฏ Requirements

  • Strong experience building/operating production systems focused on reliability, scalability, and maintainability.
  • Systems mindset: bottlenecks, failure modes, interfaces, resource usage, operability.
  • Hands-on experience with cloud infrastructure, Linux, and infrastructure automation.
  • Experience with Kubernetes and operating distributed workloads in production.
  • Strong coding skills, ideally Python or similar for backend systems.
  • Experience building internal platforms, developer tooling, or infrastructure abstractions.
  • Comfort working in ambiguous environments and taking ownership of open-ended technical problems.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’