Senior Software Engineer, AI Inference Systems
This role involves developing and optimizing high-performance AI inference systems, focusing on GPU kernel optimization, compiler infrastructure, and large-scale model deployment. You'll contribute to frameworks like vLLM, enhance inference efficiency using techniques such as speculative decoding and tensor parallelism, and build benchmarking tools for industry standards like MLPerf. The position emphasizes research, performance engineering, and collaboration across compiler, scheduling, and systems teams to advance AI inference at scale.