Sr. Site Reliability Engineer, AI Infrastructure (Starshield)

Added
12 minutes ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

site reliability ansible terraform linux python
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now โ†’

๐Ÿ“‹ Description

  • Manage GPU/CPU infrastructure deployments to Top Secret data centers
  • Manage and provide support for GPU as a service for external customers on bare metal hardware and
  • Design, validate, and productize solutions for AI clusters (100k+ GPU scale)
  • Develop automation to deploy and manage on-premise Kubernetes/AI clusters, and operating systems
  • Deploy and manage core infrastructure such as databases, monitoring and distributed storage
  • Closely collaborate with AI engineers to create highly scalable, operable, and maintainable products

๐ŸŽฏ Requirements

  • Bachelorโ€™s degree in computer science, information systems/IT, or an engineering discipline and 5+
  • 5+ year of experience with Kubernetes
  • 5+ year of experience managing Linux operating systems
  • Experience with Terraform, Ansible, or other infrastructure tools
  • Experience with containerization technologies (i.e. OCI containers, Kubernetes)
  • Experience scripting in Bash, Python, or other similar languages

๐ŸŽ Benefits

  • Base salary range $165,000.00 - $265,000.00 (Level 3)
  • Long-term incentives (company stock or long-term cash awards)
  • Discretionary bonuses
  • Employee Stock Purchase Plan
  • Comprehensive medical, vision, and dental coverage
  • 401(k) retirement plan
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’