Search

Senior Software Engineer - Fleet Automation / GPU Infrastructure

PublishedPublished: 6/14/2022
Technology

Job Description

Job Description

Senior Software Engineer – Fleet Automation / GPU Infrastructure

Location: Dallas, TX
Work Model: Hybrid, 3 days onsite
Relocation: Available
Compensation: base + bonus
Benefits: 100% company-paid benefits
Employment Type: Direct Hire

Overview

Our client is seeking a Senior Software Engineer to join its Fleet Automation team supporting large-scale HPC and GPU infrastructure.

This team builds the software, automation, and internal platforms used to provision, configure, monitor, and manage hundreds of high-performance GPU and CPU compute nodes. The environment sits at the intersection of software engineering, infrastructure, and distributed systems, with a strong focus on eliminating manual operational work through scalable automation.

This is a hands-on engineering role for someone who enjoys building backend services while also understanding how those services interact with Linux systems, physical hardware, networking, storage, and GPU infrastructure.

What You’ll Do

  • Design and build fleet automation platforms for provisioning, configuration, validation, and lifecycle management of GPU and CPU compute nodes.
  • Develop internal services and APIs that automate hardware deployment, imaging, remediation, and decommissioning.
  • Build reliable backend services using Go, C#, and/or TypeScript.
  • Design data models and persistent state for automation workflows using relational and NoSQL databases.
  • Develop and maintain CI/CD pipelines for infrastructure and configuration changes.
  • Automate hardware validation and testing across large-scale compute environments.
  • Build observability, monitoring, dashboards, and alerting using tools such as Prometheus, Grafana, Alertmanager, and ELK.
  • Partner closely with Infrastructure, Network, Operations, and Research teams to identify operational pain points and automate repeatable processes.
  • Participate in on-call rotations and support incident response, root-cause analysis, and post-incident reliability improvements.
  • Identify systemic infrastructure issues and develop engineering solutions that improve fleet reliability, efficiency, and scalability.

Required Experience

  • 5+ years of software engineering experience building production backend services, infrastructure platforms, or automation tooling.
  • Strong development experience with at least one of the following:
    • Go
    • C#
    • TypeScript
  • Experience designing APIs, backend services, and distributed or stateful systems.
  • Strong experience with relational and/or NoSQL databases.
  • Solid Linux systems knowledge, including:
    • Networking
    • Storage
    • Process management
    • System troubleshooting
    • Ubuntu and/or RHEL environments
  • Experience building and maintaining CI/CD pipelines.
  • Hands-on experience with production monitoring and observability platforms such as:
    • Prometheus
    • Grafana
    • Alertmanager
    • ELK
  • Strong troubleshooting and problem-solving skills across both software and infrastructure environments.

Preferred Experience

  • Experience supporting GPU, HPC, AI/ML, or large-scale compute infrastructure.
  • Familiarity with NVIDIA technologies such as:
    • DCGM
    • nvidia-smi
    • NVIDIA Container Toolkit
  • Experience with bare-metal provisioning and hardware lifecycle automation.
  • Exposure to event-driven architectures and messaging platforms such as Kafka.
  • Experience working with infrastructure, network, SRE, or platform engineering teams.
  • Bachelor’s degree in Computer Science, Software Engineering, or equivalent practical experience.

Why Consider This Opportunity?

  • Work directly on large-scale GPU and HPC infrastructure supporting advanced AI workloads.
  • Build automation platforms that have direct impact on infrastructure reliability and scalability.
  • Highly technical environment combining software engineering, systems engineering, and infrastructure automation.
  • Competitive base salary + bonus.
  • 100% company-paid benefits.
  • Relocation assistance available for candidates moving to Dallas.
  • Hybrid schedule with three days per week onsite.

\nCompany Description

GTN is the leader in SOW management & technical staffing, leveraging innovation to drive next-generation recruiting to Fortune 200 companies.

Company Description

GTN is the leader in SOW management & technical staffing, leveraging innovation to drive next-generation recruiting to Fortune 200 companies.

Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...