Accelerate RL rollouts by up to 50% with distribution-aware speculative decoding
Distribution-aware speculative decoding (DAS) is a novel framework that significantly alleviates the rollout bottleneck in RL post-training — delivering up to 50% speedup without touching model outputs. The rollout bottleneck…