Streaming transcription is faster, wrong, and confidently so.
Whisper-family models were built to transcribe whole utterances. Force them to stream and you trade away the future context they need to be accurate. At the bedside, where the last word often changes the first, that tradeoff is sharper than it looks.