Site Reliability Engineer, AI Infrastructure (Starshield)

Added
10 minutes ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

site reliability devops ansible terraform linux

๐Ÿ“‹ Description

  • Manage GPU/CPU infrastructure deployments to Top Secret datacenters
  • Manage and provide support for GPU as a service for external customers on bare metal hardware and
  • Design, validate, and productize solutions for AI clusters (100k+ GPU scale)
  • Develop automation to deploy and manage on-premise Kubernetes/AI clusters, and operating systems
  • Deploy and manage core infrastructure such as databases, monitoring and distributed storage
  • Closely collaborate with AI engineers to create highly scalable, operable, and maintainable products

๐ŸŽฏ Requirements

  • Bachelor's degree in computer science, information systems/IT, or an engineering discipline and 1+
  • 1+ years of professional experience with Linux operating systems
  • Experience with Terraform, Ansible, or other infrastructure tools
  • Experience with containerization technologies (i.e. OCI containers, Kubernetes)
  • Experience scripting in Bash, Python, or other similar languages
  • Development experience in Python, C++, or Go

๐ŸŽ Benefits

  • Base salary is just one part of your total rewards package at SpaceX. You may also be eligible for
  • You will also receive access to comprehensive medical, vision, and dental coverage, access to a
  • You may also accrue 3 weeks of paid vacation and will be eligible for 10 or more paid holidays per
  • Those with an active clearance will receive a 10% differential, up to an additional $20,000
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’