Senior Software Engineer, AI Inference Systems
Develop and optimize high-performance AI inference systems for large-scale models on NVIDIA GPUs. Work on inference frameworks like vLLM, optimize GPU kernels and compilers, design benchmarking methodologies, and contribute to MLPerf submissions. Collaborate across teams to improve performance in multi-GPU, multi-node, and cloud environments, with opportunities to conduct and publish research in ML systems.