Senior HPC Engineer, GPU Compute
This role involves optimizing and maintaining high-performance GPU clusters and InfiniBand networks within a large-scale AI cloud infrastructure. The engineer will work on low-level system tuning, hardware integration, and automation for fault detection across multi-GPU HPC environments. Key focus areas include performance analysis, virtualization (KVM/QEMU), and support for distributed computing workloads using technologies like Kubernetes, MPI, and NCCL.