Whisper-Flash: Acoustically Conditioned Parallel Drafting for Faster Whisper Decoding
Whisper is a widely used encoder-decoder model for speech recognition. Its encoder reads an utterance in one parallel pass, but its decoder writes the transcript one token at a time, which dominates inference time. Speculative decoding shortens such loops without changing their output: a small drafter guesses several u...