Whisper does not guarantee word-level timing
Base Whisper output commonly gives segment-level timestamps, which can be too broad for subtitles, karaoke-style highlighting, or precise editing.
Workaround
Use WhisperX alignment when individual words must map closely to the audio.