Senior ML Engineer (Token Factory)
The role involves optimizing large-scale LLM inference and fine-tuning performance across tens of thousands of GPUs, focusing on reducing latency and cost-per-token. You'll work on cutting-edge problems like speculative decoding, low-precision pipelines, and inference engine improvements, with deep involvement in transformer architectures and GPU compute efficiency. The position offers technical leadership opportunities within a globally distributed engineering team pushing the boundaries of AI infrastructure.