Model or dataset
linto-ai/whisper-timestamped avatar
linto-ai/whisper-timestamped

whisper-timestamped: word-level timestamps on top of Whisper

Multilingual Automatic Speech Recognition with word-level timestamps and confidence

2,852 stars212 forksPythonAGPL-3.0

At a glance

What is it?
whisper-timestamped wraps openai-whisper with word-level timestamps and per-word confidence, using dynamic time warping over cross-attention weights instead of a second alignment model. It is a good fit when you need subtitle-grade timing without adding a wav2vec2 pass to your pipeline.
Who is it for?
Adopt whisper-timestamped if you already run openai-whisper and need word-level timing for subtitles, forced alignment or caption QC, and you accept the AGPL-3.0 licence and the extra memory that DTW alignment adds on long files. Do not adopt it if you need a permissive licence for a closed product, or if you need speaker diarization, which the repository lists as a topic but does not document as a feature.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap whisper-timestamped fills: Whisper does not predict word timestamps

Whisper models were trained to predict approximate timestamps on speech segments, and the README states this accuracy is most of the time around one second. That is fine for a paragraph of subtitles and useless for karaoke-style highlighting, word-level search, or aligning a transcript to a specific frame in an editor. whisper-timestamped is an extension of the openai-whisper Python package, and the problem it solves is narrow: recover word-level timestamps and a confidence score per word and per segment from a model that was never trained to output them. The intended user is someone who already has openai-whisper in a pipeline and wants better timing without swapping the recognizer. The README is explicit that this is an experimental extension and that it may significantly impact performance, so it is not pitched as a drop-in production replacement for the base package. If you only need segment-level subtitles, the base Whisper output is already at the accuracy the model was trained for, and this project adds work for a precision you may not need.

How the DTW alignment works and what it costs in memory

The mechanism is dynamic time warping applied to the cross-attention weights of the Whisper decoder, an approach the README credits to a notebook by Jong Wook Kim. During decoding, the attention the decoder pays to each audio frame is recorded, and DTW finds the path that maps predicted tokens onto that attention matrix. The README lists three additions over that notebook: more accurate start and end estimation, confidence scores assigned to each word, and, when beam search is not in use, no extra inference steps, because alignment runs on the fly after each speech segment is decoded. The phrase about beam search matters. Turn beam search on and the on-the-fly property no longer holds, so the cost profile changes. The README also states that special care was taken with memory so that long files can be processed with little additional memory compared to regular Whisper use. That claim is worth treating as a design goal rather than a guarantee: DTW over an attention matrix scales with audio length and token count, and the README does not publish a table of measured memory against file duration.

Installing whisper-timestamped with pip and running a first transcription

The README lists python3 (3.7 or higher, with 3.9 recommended) and ffmpeg as requirements, and gives pip as the primary install path. The package installs three dependencies from requirements.txt: Cython, dtw-python and openai-whisper. Run the install, then transcribe a file. The README's example output section shows the structure you should expect: segments, each with words, and each word carrying start, end and confidence fields.

bash
pip3 install whisper-timestamped

The same package can be installed from a clone, which is what the Docker build does. The README gives this alternative for source installs.

bash
git clone https://github.com/linto-ai/whisper-timestamped
cd whisper-timestamped/
python3 setup.py install

Optional extras are installed separately. matplotlib is needed for the alignment plot, onnxruntime and torchaudio for the VAD option, and transformers for finetuned Whisper models from the Hugging Face Hub.

bash
pip3 install matplotlib
pip3 install onnxruntime torchaudio
pip3 install transformers

The repository also ships a Dockerfile that the README says builds an image of about 9GB, plus a Dockerfile.cpu for machines without a GPU. The Dockerfile installs the extras through setup.py targets and sets bash as the entrypoint.

VAD, disfluencies and the options that change accuracy

The README devotes a section to options that may improve results, and the most consequential is voice activity detection before Whisper runs. The stated reason is hallucination: Whisper was trained on data containing trailing phrases, so on pure silence it can emit text like "Thanks you for watching!". Running VAD first removes those regions before the model sees them. Several methods are available, with silero as the default and auditok and auditok:v3.1 as alternatives. This is a real design choice with a trade-off: VAD that is too aggressive clips quiet speech at the start of a word, and since word boundaries are exactly what this project is trying to measure, a false negative in VAD propagates into the timestamps. The README also covers detecting disfluencies, and when the language is not specified, the language probabilities appear among the outputs, which is useful for routing multilingual audio without a separate classifier. None of these options are on by default beyond the silero VAD choice, so a default run and a tuned run can produce noticeably different word boundaries on the same file.

Where the cross-attention approach breaks down

The README is unusually direct about competing methods, and its criticism of the timestamp-token approach applies to its own limits too. Whisper models were not trained to output meaningful timestamps after each word; they tend to predict timestamps only after a number of words, typically at the end of a sentence, and the probability distribution outside that condition can be inaccurate. The README says these methods can produce results that are totally out of sync on some periods, especially with jingle music, and that Whisper's timestamp precision tends to be rounded to one second. The DTW approach avoids the token-probability problem but does not escape the underlying model. If Whisper mis-transcribes a word or drops a filler, there is no token to align, and the word simply does not appear in the output with a timestamp. Music, overlapping speakers and heavy background noise remain hard. The README's own disclaimer calls the extension experimental and warns of significant performance impact, so treating its timestamps as ground truth for legal or medical transcription would be a mistake.

whisper-timestamped vs whisperX and the wav2vec2 route

The README names whisperX as the relevant alternative and explains the difference in approach rather than dismissing it. whisperX recovers word-level timestamps by running wav2vec models that predict characters, which means one alignment model per language. The README lists the drawbacks it sees: the per-language model requirement does not scale with Whisper's multilingual coverage, an extra neural network consumes memory, and the transcription has to be normalized to match the wav2vec character set, which means language-dependent conversions such as turning "2" into "two" or "%" into "percent". It also notes that disfluencies Whisper usually removes are a problem for character-level alignment. The counter-argument is implicit: wav2vec2 alignment is a trained acoustic model doing the alignment task, while DTW over attention weights is a post-hoc interpretation of what the decoder already computed. Which is more accurate on your audio is not something the README settles, and it is the first thing to measure before choosing.

Licence, maintenance and upgrade cost

The licence is AGPL-3.0, which is the single biggest adoption constraint here. If you run whisper-timestamped as part of a network service, the AGPL's source-disclosure obligations reach users who interact with that service over a network. For internal tooling this is usually a non-issue; for a commercial product with a closed codebase it is a decision for your legal team, not a technical one. On maintenance, the last push to the default branch was on 2026-08-17, and the most recent release listed is v1.15.9 from 2025-09-09, with the two releases before that in November 2024. The repository is not archived. Upgrades are cheap in dependency terms because requirements.txt pins only Cython, dtw-python and openai-whisper, and setup.py asserts that requirements.txt matches install_requires, so a mismatch fails the build rather than silently diverging. The real upgrade cost is behavioural: because the package is tied to openai-whisper and meant to be compatible with any version of it, changes in Whisper's attention internals or decoding loop are the things most likely to move your timestamps between versions.

Editorial conclusion

Adopt whisper-timestamped if you already run openai-whisper and need word-level timing for subtitles, forced alignment or caption QC, and you accept the AGPL-3.0 licence and the extra memory that DTW alignment adds on long files. Do not adopt it if you need a permissive licence for a closed product, or if you need speaker diarization, which the repository lists as a topic but does not document as a feature. Before committing, run it on one of your own audio files and inspect the words array in the JSON output against a segment you have transcribed by hand.

Frequently asked questions

Can I transcribe audio with timestamps using whisper-timestamped?

Yes. The package adds word-level timestamps and confidence scores to Whisper transcriptions, and the README shows JSON output containing segments with words that each carry start, end and confidence values.

How can I transcribe a video with timestamps?

The README lists ffmpeg as a requirement alongside python3, so video audio can be extracted first and then passed to the transcribe function, which returns word timestamps in the JSON output.

How can I add timestamps to audio?

Install with pip3 install whisper-timestamped, then call the transcribe function and read the words array in the returned result, where each word carries its own start and end values.

What is timestamping in transcription?

It is the assignment of time positions to recognized speech. Whisper itself predicts approximate timestamps on speech segments, most of the time with one-second accuracy, and whisper-timestamped extends that to individual words with a confidence score each.

Official sources

  1. Issues
  2. License: AGPL-3.0
  3. linto-ai/whisper-timestamped on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/linto-ai-whisper-timestamped.svg)](https://hysenlabs.com/projects/linto-ai-whisper-timestamped)