Search

Network Engineer (AI GPU Cluster Operations)

PublishedPublished: 6/14/2022
Technology

Job Description

Position Overview

\n

The AI GPU Cluster Operations – Network Engineer is responsible for the operations, maintenance and troubleshooting of high-performance networking infrastructure supporting large-scale AI GPU computing clusters.

\n

The role focuses on InfiniBand, RoCE and data center IP networks, while also requiring a strong understanding of GPU servers, Linux and the overall AI cluster architecture.

\n

The engineer will monitor network and cluster health, troubleshoot connectivity and performance issues, maintain high-speed network infrastructure, support production incidents, implement approved network changes and work with internal engineering teams and OEM vendors to drive technical issues through resolution.

\n

Key Responsibilities

\n

• Operate and maintain high-performance GPU cluster networks, including InfiniBand and RoCE, as well as standard data center IP networks.

\n

• Manage and troubleshoot NVIDIA/Mellanox InfiniBand switches, high-speed Ethernet switches, NICs/HCAs, optical transceivers, cables and related network infrastructure.

\n

• Monitor network health and identify link degradation, link flapping, congestion, packet loss, bandwidth or latency issues.

\n

• Perform InfiniBand health checks and diagnostics using tools such as ibstat, ibdiagnet and other NVIDIA/Mellanox diagnostic utilities.

\n

• Support InfiniBand Subnet Management, partition configuration, routing, link-health monitoring and performance-counter analysis.

\n

• Troubleshoot RoCE environments, including connectivity, congestion and lossless Ethernet-related issues.

\n

• Support data center IP network operations, including BGP, OSPF, VLAN, MLAG and TCP/IP.

\n

• Troubleshoot GPU node connectivity and determine whether infrastructure issues originate from the server, NIC/HCA, optics, cabling, switch or network fabric.

\n

• Work with GPU Cluster Operations Engineers to troubleshoot GPU-to-GPU and node-to-node communication issues, including network-related NCCL failures and performance degradation.

\n

• Monitor network and infrastructure health using Prometheus, SNMP, Grafana and related monitoring platforms.

\n

• Develop and maintain network monitoring, alerting and diagnostic capabilities.

\n

• Perform approved network configuration changes, maintenance and upgrades according to established change-management procedures and Standard Operating Procedures (SOPs).

\n

• Support network capacity planning and expansion of GPU cluster infrastructure.

\n

• Use Python, Shell, Ansible and/or Terraform to automate network configuration, health checks, monitoring and repetitive operational tasks.

\n

• Participate in infrastructure incident response, troubleshooting, post-incident reviews and Root Cause Analysis (RCA).

\n

• Maintain network diagrams, configuration documentation, SOPs, troubleshooting guides and operational records.

\n

• Coordinate with NVIDIA/Mellanox, server OEMs, network vendors and other technical support teams to resolve hardware and network issues.

\n

• Support hardware replacement and RMA activities for switches, NICs/HCAs, optical transceivers and other network components.

\n

Qualifications

\n

• Degree or relevant educational background in Computer Science, Information Technology, Networking, Telecommunications, Engineering or a related field.

\n

• 3+ years of hands-on network operations, data center networking or infrastructure experience.

\n

• Hands-on experience operating InfiniBand and/or RoCE networks in GPU, AI, HPC or large-scale data center environments.

\n

• Experience with NVIDIA/Mellanox InfiniBand or high-speed Ethernet networking equipment is strongly preferred.

\n

• Understanding of InfiniBand architecture, including Subnet Management, partitions, routing, link states and performance counters.

\n

• Strong understanding of TCP/IP and data center networking.

\n

• Working knowledge of BGP, OSPF, VLAN and MLAG.

\n

• Ability to troubleshoot high-speed network connectivity, link degradation, congestion, packet loss and performance issues.

\n

• Hands-on ability to troubleshoot switches, NICs/HCAs, optical transceivers, cables and switch ports.

\n

• Familiarity with InfiniBand diagnostic tools such as ibstat and ibdiagnet.

\n

• Familiarity with Linux and command-line troubleshooting.

\n

• Experience with monitoring platforms and protocols such as Prometheus, SNMP and Grafana.

\n

• Ability to write basic Python and/or Shell scripts for troubleshooting and automation.

\n

• Experience with Ansible and/or Terraform is preferred.

\n

• Good understanding of GPU cluster architecture and the relationship between GPU servers, NICs/HCAs and high-performance network fabrics.

\n

• Strong troubleshooting, documentation and Root Cause Analysis (RCA) capabilities.

\n

• Strong sense of ownership and ability to work effectively during production incidents.

\n

Preferred Qualifications

\n

• Experience supporting large-scale 1,000+ GPU clusters; experience with 10,000+ GPU environments is a strong plus.

\n

• Experience supporting NVIDIA H100/H200/B200 or AMD MI300X/MI355X GPU infrastructure.

\n

• Experience operating 100G/200G/400G or higher-speed InfiniBand or Ethernet networks.

\n

• Experience troubleshooting NCCL and GPU communication issues from the network/fabric perspective.

\n

• Experience with NVIDIA/Mellanox network management and diagnostic tools.

\n

• Experience with network automation using Python, Shell, Ansible or Terraform.

\n

• CCNP, CCIE or other relevant networking certifications are a plus.

\n

• Experience working within ITIL-based incident, problem and change-management processes.

Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...