Search

GPU Infrastructure Engineer

PublishedPublished: 6/14/2022
Technology

Job Description

Key Responsibilities

\n

    \n
  • Design, deploy, maintain, and troubleshoot enterprise-scale Linux and Windows server infrastructure.
  • \n

  • Support hyperscale GPU and compute environments used for AI training and inference workloads.
  • \n

  • Troubleshoot NVIDIA and AMD GPU platforms, including GPU memory errors, driver issues, CUDA/ROCm failures, PCIe problems, and system-level hardware faults.
  • \n

  • Work with GPU technologies and related OCP-based hardware to optimize system performance.
  • \n

  • Diagnose server, rack, power, networking, PCIe, GPU, NIC, and operating-system issues.
  • \n

  • Utilize Linux kernel tools, NVIDIA/AMD diagnostic utilities, system logs, and custom troubleshooting tools to identify and resolve infrastructure problems.
  • \n

  • Develop automation and infrastructure tooling using Python, Bash, Ansible, Chef, and related technologies.
  • \n

  • Improve operational tools and automation used across large server fleets.
  • \n

  • Support VMware and Hyper-V virtualization environments where required.
  • \n

  • Administer and troubleshoot networking infrastructure, including switches, firewalls, and data center connectivity.
  • \n

  • Monitor infrastructure health and performance using tools such as Grafana, InfluxDB, Telegraf, Nagios, and related monitoring platforms.
  • \n

  • Participate in on-call rotations and provide critical incident response for production infrastructure.
  • \n

  • Identify systemic infrastructure issues and develop scalable solutions to prevent recurrence.
  • \n

  • Collaborate with engineering and infrastructure teams to improve reliability, availability, and operational efficiency.
  • \n

\n

. Required Qualifications

\n

    \n
  • 10+ years of experience in systems, infrastructure, production engineering, or related roles.
  • \n

  • Strong Linux administration and troubleshooting experience, preferably with RHEL and Ubuntu.
  • \n

  • Hands-on experience with enterprise server and data center infrastructure.
  • \n

  • Experience troubleshooting GPU-based infrastructure and high-performance compute environments.
  • \n

  • Strong understanding of server hardware, PCIe, networking, storage, and operating-system diagnostics.
  • \n

  • Programming/scripting experience with Python and Bash.
  • \n

  • Experience with infrastructure automation tools such as Ansible or Chef.
  • \n

  • Strong troubleshooting and root-cause-analysis skills.
  • \n

  • Experience working in large-scale production environments.
  • \n

  • Excellent communication and cross-functional collaboration skills.
  • \n

\n

Ideal Candidate Profile

\n

The ideal candidate is a hands-on infrastructure professional who can operate at both the hardware and software layers, from diagnosing GPU/server failures and Linux kernel issues to developing automation that improves reliability across large-scale infrastructure. Experience supporting hyperscale AI/GPU environments, combined with strong systems engineering and production troubleshooting skills, is highly valued.

\n


Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...