Added
8 hours ago
Type
Full time
Salary
Salary not provided

Related skills

datadog terraform github actions linux aws

πŸ“‹ Description

  • Design, build, and improve platform systems for ML training, deployment, and serving.
  • Build scalable infra and tools to improve ML workload reliability, efficiency, and cost.
  • Develop automation workflows and tools to reduce complexity for researchers and engineers.
  • Architect and maintain systems for deployment, monitoring, and operation across environments.
  • Improve scheduling, monitoring, debugging, and resource management for GPU and cloud infra.
  • Improve observability, automation, reliability, and developer usability.

🎯 Requirements

  • Strong experience building/operating production systems with reliability and scale.
  • Systems mindset to reason about bottlenecks, failures, interfaces, and resources.
  • Hands-on cloud infra, Linux, and infrastructure automation experience.
  • Experience operating distributed systems in production, incl. Kubernetes.
  • Strong Python or backend-language programming skills.
  • Experience building internal platforms and developer tooling.

🎁 Benefits

  • Fully remote work environment with flexibility across Europe.
  • Opportunity to work on advanced AI infrastructure supporting large-scale enterprise applications.
  • High ownership role with significant influence over platform architecture and technical direction.
  • Collaborative environment with close interaction between engineering, research, and product teams.
  • Opportunity to solve complex challenges involving machine learning systems, automation, and distributed infrastructure.
  • Professional growth opportunities within a fast-moving AI-focused organization.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’