Network Engineer (AI GPU Cluster Operations)
Job Description
Position Overview
\n
The AI GPU Cluster Operations – Network Engineer is responsible for the operations, maintenance and troubleshooting of high-performance networking infrastructure supporting large-scale AI GPU computing clusters.
\n
The role focuses on InfiniBand, RoCE and data center IP networks, while also requiring a strong understanding of GPU servers, Linux and the overall AI cluster architecture.
\n
The engineer will monitor network and cluster health, troubleshoot connectivity and performance issues, maintain high-speed network infrastructure, support production incidents, implement approved network changes and work with internal engineering teams and OEM vendors to drive technical issues through resolution.
\n
Key Responsibilities
\n
• Operate and maintain high-performance GPU cluster networks, including InfiniBand and RoCE, as well as standard data center IP networks.
\n
• Manage and troubleshoot NVIDIA/Mellanox InfiniBand switches, high-speed Ethernet switches, NICs/HCAs, optical transceivers, cables and related network infrastructure.
\n
• Monitor network health and identify link degradation, link flapping, congestion, packet loss, bandwidth or latency issues.
\n
• Perform InfiniBand health checks and diagnostics using tools such as ibstat, ibdiagnet and other NVIDIA/Mellanox diagnostic utilities.
\n
• Support InfiniBand Subnet Management, partition configuration, routing, link-health monitoring and performance-counter analysis.
\n
• Troubleshoot RoCE environments, including connectivity, congestion and lossless Ethernet-related issues.
\n
• Support data center IP network operations, including BGP, OSPF, VLAN, MLAG and TCP/IP.
\n
• Troubleshoot GPU node connectivity and determine whether infrastructure issues originate from the server, NIC/HCA, optics, cabling, switch or network fabric.
\n
• Work with GPU Cluster Operations Engineers to troubleshoot GPU-to-GPU and node-to-node communication issues, including network-related NCCL failures and performance degradation.
\n
• Monitor network and infrastructure health using Prometheus, SNMP, Grafana and related monitoring platforms.
\n
• Develop and maintain network monitoring, alerting and diagnostic capabilities.
\n
• Perform approved network configuration changes, maintenance and upgrades according to established change-management procedures and Standard Operating Procedures (SOPs).
\n
• Support network capacity planning and expansion of GPU cluster infrastructure.
\n
• Use Python, Shell, Ansible and/or Terraform to automate network configuration, health checks, monitoring and repetitive operational tasks.
\n
• Participate in infrastructure incident response, troubleshooting, post-incident reviews and Root Cause Analysis (RCA).
\n
• Maintain network diagrams, configuration documentation, SOPs, troubleshooting guides and operational records.
\n
• Coordinate with NVIDIA/Mellanox, server OEMs, network vendors and other technical support teams to resolve hardware and network issues.
\n
• Support hardware replacement and RMA activities for switches, NICs/HCAs, optical transceivers and other network components.
\n
Qualifications
\n
• Degree or relevant educational background in Computer Science, Information Technology, Networking, Telecommunications, Engineering or a related field.
\n
• 3+ years of hands-on network operations, data center networking or infrastructure experience.
\n
• Hands-on experience operating InfiniBand and/or RoCE networks in GPU, AI, HPC or large-scale data center environments.
\n
• Experience with NVIDIA/Mellanox InfiniBand or high-speed Ethernet networking equipment is strongly preferred.
\n
• Understanding of InfiniBand architecture, including Subnet Management, partitions, routing, link states and performance counters.
\n
• Strong understanding of TCP/IP and data center networking.
\n
• Working knowledge of BGP, OSPF, VLAN and MLAG.
\n
• Ability to troubleshoot high-speed network connectivity, link degradation, congestion, packet loss and performance issues.
\n
• Hands-on ability to troubleshoot switches, NICs/HCAs, optical transceivers, cables and switch ports.
\n
• Familiarity with InfiniBand diagnostic tools such as ibstat and ibdiagnet.
\n
• Familiarity with Linux and command-line troubleshooting.
\n
• Experience with monitoring platforms and protocols such as Prometheus, SNMP and Grafana.
\n
• Ability to write basic Python and/or Shell scripts for troubleshooting and automation.
\n
• Experience with Ansible and/or Terraform is preferred.
\n
• Good understanding of GPU cluster architecture and the relationship between GPU servers, NICs/HCAs and high-performance network fabrics.
\n
• Strong troubleshooting, documentation and Root Cause Analysis (RCA) capabilities.
\n
• Strong sense of ownership and ability to work effectively during production incidents.
\n
Preferred Qualifications
\n
• Experience supporting large-scale 1,000+ GPU clusters; experience with 10,000+ GPU environments is a strong plus.
\n
• Experience supporting NVIDIA H100/H200/B200 or AMD MI300X/MI355X GPU infrastructure.
\n
• Experience operating 100G/200G/400G or higher-speed InfiniBand or Ethernet networks.
\n
• Experience troubleshooting NCCL and GPU communication issues from the network/fabric perspective.
\n
• Experience with NVIDIA/Mellanox network management and diagnostic tools.
\n
• Experience with network automation using Python, Shell, Ansible or Terraform.
\n
• CCNP, CCIE or other relevant networking certifications are a plus.
\n
• Experience working within ITIL-based incident, problem and change-management processes.
