Added
13 days ago
Type
Full time
Salary
Salary not provided

Related skills

docker ansible terraform grafana prometheus
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now β†’

πŸ“‹ Description

  • Design, scale, and maintain high-performance MLOps and LLMOps infrastructure for an AI-infused
  • Collaborate with Data Scientists, ML Engineers, and Cloud Infra teams to streamline training
  • Ensure GPU utilisation, reliability, and cost-efficiency across cloud environments

🎯 Requirements

  • Hands-on production experience in DevOps, SRE, or Platform Engineering for AI/ML infra
  • Proven track record deploying and operationalising ML models and LLMs in cloud-native production
  • Experience managing GPU infrastructure and HPC environments
  • Advanced proficiency in Kubernetes, Docker, Helm, KubeFlow; service meshes (e.g., Istio)
  • Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins
  • Experience with vLLM, Ray, MLflow, LangChain / LangSmith, DeepSpeed, or Hugging Face TGI

🎁 Benefits

  • DEIB focused culture and opportunities to grow within a global team
  • Competitive benefits and career development
  • Work on cutting-edge AI/ML infrastructure and platform engineering
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’