Network Operations Engineer, AI Networking

Added
17 days ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

terraform grafana prometheus python rest apis
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now →

📋 Description

  • Own operational health of production AI network infra across 1P/3P data centers.
  • Monitor, troubleshoot, and resolve network incidents to meet SLOs.
  • Operate and maintain large-scale Ethernet fabrics for GPU compute, storage, and management networks.
  • Execute production network changes and capacity expansions with minimal impact.
  • Manage hardware lifecycle: switches, optics, RMAs, software upgrades, maintenance.
  • Support AI cluster deployments, data center expansions, and migrations.

🎯 Requirements

  • Bachelor’s degree in CS, Network Eng, or related field, or equivalent.
  • 5+ years in large-scale data center/cloud/AI/HPC network infra.
  • Experience supporting production networks with high availability.
  • Hands-on with Cisco NX-OS, Arista EOS, NVIDIA Spectrum, or Juniper JunOS.
  • L2/L3 networking knowledge: BGP/OSPF/ECMP/VRFs/VLANs.
  • Troubleshoot fiber optics, transceivers, DAC/AOC cables, and high-speed links.

🎁 Benefits

  • 24x7 on-call rotation for mission-critical AI infra.
  • Support time-sensitive incidents and maintenance with minimal impact.
  • Up to 30% travel to data centers for turnups.
  • Work with cutting-edge GPU AI infrastructure.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →