Job Description
Key Responsibilities
\n
- \n
- Design, deploy, maintain, and troubleshoot enterprise-scale Linux and Windows server infrastructure.
- Support hyperscale GPU and compute environments used for AI training and inference workloads.
- Troubleshoot NVIDIA and AMD GPU platforms, including GPU memory errors, driver issues, CUDA/ROCm failures, PCIe problems, and system-level hardware faults.
- Work with GPU technologies and related OCP-based hardware to optimize system performance.
- Diagnose server, rack, power, networking, PCIe, GPU, NIC, and operating-system issues.
- Utilize Linux kernel tools, NVIDIA/AMD diagnostic utilities, system logs, and custom troubleshooting tools to identify and resolve infrastructure problems.
- Develop automation and infrastructure tooling using Python, Bash, Ansible, Chef, and related technologies.
- Improve operational tools and automation used across large server fleets.
- Support VMware and Hyper-V virtualization environments where required.
- Administer and troubleshoot networking infrastructure, including switches, firewalls, and data center connectivity.
- Monitor infrastructure health and performance using tools such as Grafana, InfluxDB, Telegraf, Nagios, and related monitoring platforms.
- Participate in on-call rotations and provide critical incident response for production infrastructure.
- Identify systemic infrastructure issues and develop scalable solutions to prevent recurrence.
- Collaborate with engineering and infrastructure teams to improve reliability, availability, and operational efficiency.
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
. Required Qualifications
\n
- \n
- 10+ years of experience in systems, infrastructure, production engineering, or related roles.
- Strong Linux administration and troubleshooting experience, preferably with RHEL and Ubuntu.
- Hands-on experience with enterprise server and data center infrastructure.
- Experience troubleshooting GPU-based infrastructure and high-performance compute environments.
- Strong understanding of server hardware, PCIe, networking, storage, and operating-system diagnostics.
- Programming/scripting experience with Python and Bash.
- Experience with infrastructure automation tools such as Ansible or Chef.
- Strong troubleshooting and root-cause-analysis skills.
- Experience working in large-scale production environments.
- Excellent communication and cross-functional collaboration skills.
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
Ideal Candidate Profile
\n
The ideal candidate is a hands-on infrastructure professional who can operate at both the hardware and software layers, from diagnosing GPU/server failures and Linux kernel issues to developing automation that improves reliability across large-scale infrastructure. Experience supporting hyperscale AI/GPU environments, combined with strong systems engineering and production troubleshooting skills, is highly valued.
\n
