##article.return## Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding Download Download PDF