Added
1 day ago
Type
Full time
Salary
Salary not provided

Related skills

datadog terraform github actions linux python

πŸ“‹ Description

  • Design and improve platform systems for ML training, evaluation, deployment, and production serving.
  • Build scalable infra and tools to improve reliability, efficiency, and cost of ML workloads.
  • Develop automation workflows and internal tools to reduce operational complexity for researchers and engineers.
  • Architect and maintain systems enabling deployment, monitoring, and operation of ML models across research and product environments.
  • Improve scheduling, monitoring, debugging, and resource management for GPU-based and cloud infra.
  • Enhance observability, automation, reliability, developer experience, and platform usability.
  • Create abstractions and tools enabling engineers to work more effectively with ML systems.
  • Collaborate with research and product teams to turn challenges into scalable platform capabilities.
  • Contribute to architectural decisions, technical strategy, and long-term platform evolution.
  • Take ownership of open-ended engineering challenges, balancing scalability, simplicity, and reliability.

🎯 Requirements

  • Strong experience building/operating production systems focused on reliability, scalability, performance, and maintainability.
  • Systems mindset to reason about bottlenecks, failure scenarios, interfaces, resource use, and long-term needs.
  • Hands-on with cloud infrastructure, Linux, and infrastructure automation.
  • Experience operating distributed production systems, including Kubernetes-based workloads.
  • Strong Python or backend programming skills.
  • Experience building internal platforms, tooling, or developer abstractions used by engineers.
  • Understanding of ML infrastructure, model serving, or data-intensive workloads.
  • Experience with GPU-based systems or large-scale computing resources.
  • Familiarity with observability, monitoring, and debugging for distributed systems.
  • Knowledge of Terraform, Datadog, GitHub Actions, or similar tools.
  • Ability to work in ambiguous environments, take ownership, and solve complex problems independently.
  • Pragmatic engineering approach, delivering valuable solutions without unnecessary complexity.

🎁 Benefits

  • Fully remote work environment with flexibility across Europe.
  • Opportunity to work on advanced AI infrastructure powering large-scale enterprise applications.
  • High ownership role with influence over platform architecture and technical direction.
  • Collaborative environment with close interaction between engineering, research, and product teams.
  • Opportunity to solve complex challenges involving ML systems, automation, and distributed infrastructure.
  • Professional growth opportunities within a fast-moving AI-focused organization.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’