Staff HPC Systems Software Engineer
This role involves defining the technical direction for a core HPC systems domain within a cloud-native AI infrastructure platform. You'll architect and evolve Slurm-based HPC services, drive cross-team integration with Kubernetes-adjacent systems and infrastructure APIs, and establish reusable patterns for automation, reliability, and operational supportability. The position requires deep systems expertise in HPC, GPU scheduling, and production-scale software engineering to ensure robust, maintainable platform capabilities.