Site Reliability Engineer, AI Infrastructure (Starshield)

Added
2 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

ansible terraform linux bash python

πŸ“‹ Description

  • Manage GPU/CPU infrastructure deployments to Top Secret datacenters
  • Design, validate, and productize solutions for AI clusters (100k+ GPU scale)
  • Develop automation to deploy and manage on-premise Kubernetes/AI clusters, and operating systems
  • Deploy and manage core infrastructure such as databases, monitoring and distributed storage
  • Closely collaborate with AI engineers to create highly scalable, operable, and maintainable products
  • Engage in and improve the whole lifecycle of services -- from inception and design, through

🎯 Requirements

  • Bachelor's degree in computer science, information systems/IT, or an engineering discipline and 1+
  • 1+ years of professional experience with Linux operating systems
  • Experience with Terraform, Ansible, or other infrastructure tools
  • Experience with containerization technologies (i.e. OCI containers, Kubernetes)
  • Experience scripting in Bash, Python, or other similar languages
  • Development experience in Python, C++, or Go

🎁 Benefits

  • Base salary range of $125,000 - $200,000 per year depending on level
  • Long-term incentives in the form of company stock or long-term cash awards
  • Potential discretionary bonuses and the ability to purchase additional stock at a discount through
  • Comprehensive medical, vision, and dental coverage
  • Access to a 401(k) retirement plan
  • Short and long-term disability insurance
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’