AI Infrastructure Engineer
Location: Remote — must be available during Eastern Time business hours Engagement: 6–12-month contract
Industry: Banking; Financial Services
About the Role: We are seeking a hands-on AI Infrastructure Engineer to help stand up and operate a new enterprise GPU compute environment. The infrastructure includes an 18-node NVIDIA HGX cluster with 144 GPUs, supported by NVIDIA Infini Band and high-speed Ethernet switching. This is a hybrid infrastructure role spanning three areas: GPU fabric and high-speed networking Broader network operations and datacenter connectivityGPU platform, orchestration, and workload enablement The ideal candidate can operate across the full stack—from physical switch ports and fabric telemetry through GPU drivers, workload schedulers, containers, and production AI workloads. This is not a pure network engineering or pure platform engineering position.
Key Responsibilities: GPU Fabric and High-Speed Networking Manage and operate an Infini Band XDR / Quantum-3400 rail-aligned fabric topology supporting an 18-node, 144-GPU NVIDIA HGX cluster. Administer NVIDIA UFM and NetQ for subnet management, telemetry, monitoring, and firmware management. Support 800G optics, troubleshoot link-flap issues, and manage congestion isolation. Run NCCL and HPL validation and performance regression testing to confirm cluster health and throughput following infrastructure changes. Operate Cumulus SN5600 and Arista border switches across the Ethernet side of the environment. Diagnose performance and connectivity issues across multi-node GPU workloads. Network Operations Manage and support large-scale IP networking technologies and infrastructure. Monitor network health across on-premises and cloud environments. Work with peering and datacenter interconnect technologies, including: Private Network Interconnects (PNIs)Transit Internet Exchanges Passive DWDMWave circuits Improve change-management processes and day-to-day network operations. Document technical standards, operating procedures, and infrastructure best practices. Develop workflow enhancements that improve reliability and operational efficiency. GPU Platform and Workload Enablement Manage the lifecycle of NVIDIA AI Enterprise and the NVIDIA GPU Operator. Maintain driver, firmware, CUDA, and platform compatibility across the cluster. Administer Run, including quotas, project structures, resource allocation, and tenant onboarding. Configure and support Slurm across dedicated worker-node pools for scheduled batch workloads. Manage container images, runtime dependencies, and model-serving components across production and pre-production environments. Partner with the DNA MLOps team to onboard the first production workloads onto the new cluster. Support engineering and data science teams as adoption of the GPU platform expands.
Required Qualifications: Hands-on experience supporting NVIDIA Infini Band, preferably the Quantum series, within an HPC or GPU cluster environment. Experience with high-speed Ethernet switching in a multi-node GPU environment. Familiarity with NVIDIA HGX architecture and rail-aligned GPU networking topologies. Experience using NCCL, HPL, or comparable tools to validate GPU cluster health, connectivity, and performance. Working knowledge of the NVIDIA GPU Operator, GPU driver and firmware lifecycle management, and compatibility planning. Experience deploying and supporting container-based workloads in an HPC or AI environment. Experience with at least one workload orchestration platform, such as Slurm, Run, or an equivalent scheduler. Strong knowledge of large-scale IP networking and datacenter peering/interconnect technologies. Experience working within structured change-management and production operations processes. Ability to troubleshoot across physical infrastructure, network fabric, operating systems, containers, orchestration platforms, and application workloads.
Preferred Qualifications: Experience administering NVIDIA UFM, NetQ, or equivalent Infini Band fabric-management tools. Experience with new GPU cluster bring-up or greenfield AI infrastructure deployments. Exposure to MLOps pipelines, production AI workloads, and model-serving infrastructure. Familiarity with DWDM and wave circuit technologies in a datacenter interconnect environment. Experience supporting enterprise GPU infrastructure in a regulated or highly controlled environment.
Ideal Candidate Profile: The strongest candidate will combine deep GPU-cluster networking knowledge with practical platform engineering experience. They should be comfortable moving between Infini Band troubleshooting, Ethernet and datacenter operations, GPU software lifecycle management, workload scheduling, and production AI enablement.

Also on the board Same function, level within a rung

Level

Lead

Location

Denver, CO

Occupation

Network and Computer Systems Administrators

Industry

Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services

Posted

2 days ago

Apply for this role →
Artificial Intelligence Engineer at tiger advisory | Johnson Jobs