Model or dataset
octimot/StoryToolkitAI avatar
octimot/StoryToolkitAI

StoryToolkitAI: a feature list with version gates, and a release from February 2025

An editing tool that uses AI to transcribe, understand content and search for anything in your footage, integrated with ChatGPT and other AI models

1,021 stars85 forksPythonGPL-3.0

At a glance

What is it?
A local transcription and indexing tool for filmmakers that turns speech into searchable transcripts, groups them into stories, and hands the result to your editor as EDL, XML or Fountain. Everything runs on your own machine. The parts worth reading closely are the version gates on the feature list, the three dependencies installed straight from git, and the two builds the project keeps promising to merge.
Who is it for?
StoryToolkitAI is a genuine local tool rather than a wrapper, and the privacy story is the strongest thing about it: transcription, translation and search stay on the machine, and only the Assistant needs a provider. The practical question is which build you get.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 70 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The newest release is from February 2025 and the README says so

Three releases are visible. v0.23.1 on 2024-01-24, described as stories and speakers alpha, v0.23.2 on 2024-01-26 as stories and speakers beta, and v0.25.1 on 2025-02-20, described as projects and more LLMs.

The default branch was pushed on 2026-07-28. So the last tagged build is a long way behind the code that a person reading the repository is actually looking at.

The documentation does not hide this. It says that to download the latest standalone release you should use the releases page, and then immediately adds that the standalone releases will most likely always be behind the git version, so if you are comfortable with the terminal you are recommended to install from source. Detailed steps live in INSTALLATION.md.

That is a candid arrangement, but it inverts the usual trust relationship: the binary a user downloads is the older artifact, and the newest behaviour is only reachable by building it. Anyone evaluating a feature from a release binary may not have that feature at all.

The feature list doubles as a version table

Several entries in the key features list carry a version gate in parentheses, which turns a checkbox list into a changelog.

The story editor, which writes screenplays containing your transcripts and exports them for editing as EDL, XML or Fountain, is marked v. 0.20.1 and later. Translating transcripts to other languages with OpenAI GPT is v. 0.22.0 and later. Asking AI to create stories and selections from your footage is also v. 0.22.0. Automatic speaker detection in transcripts is v. 0.23.0 and later.

The newest release is v0.25.1, so the highest gate on the list, speaker detection, is two minor versions below the last tag. Read the other way, everything gated on the list exists in the newest release, and anything you read about on the branch may be newer than what the gates describe.

The rest of the list is ungated and undated: full video indexing and search, free automatic transcription and translation on your local machine, compatibility with OpenAI, ollama, vLLM and LM Studio, transcript groups, question detection, multi-format export including SRT, TXT, AVID DS and a Fusion Text node, and import of existing SRT files.

A separate file, FEATURES.md, is pointed to for the detailed version information.

Some features live in the other build until a future release

There are two versions of the tool, and the README draws the line between them in a single sentence: some of the listed features are only available in the non-standalone version, and they will be available in the standalone version in the next release.

No feature names are attached to that sentence, so a reader has no way to tell which capabilities are missing from the build they just downloaded, only that something is. The next release has not arrived, since the newest tag is from February 2025.

The distinction is not cosmetic. The standalone build is what the releases page serves, and it is the one a user installs without touching a terminal, so it is the one that determines whether the tool can do the thing you installed it for.

The project does publish a feature list and a changelog, and the Resolve integration section is written with the same ambiguity, mixing items that sound like they need the editor installed with items that are internal to the tool.

A joke item is ticked and three real ones are not

The planned features section is four lines, and the last one is marked done.

The first three are unchecked and they are the substantive ones: automatic topic classification to help you discover ideas in your transcripts, integration with other AI tools, and integration with other software and standalone players.

The fourth is checked and reads: plus more flashy features as clickbait to unrealistically raise expectations and destroy competition.

It is the only item in that list that is complete, which tells you something about the roadmap's actual state. The three unfinished items are all about reach, extending the tool beyond the Resolve workflow it is built around, and the one that is finished is a joke about expectations.

Given that the neighbouring sections are careful about version gates, the joke is the one line in the file where the author is speaking rather than documenting, and it sits exactly where a reader will find it when deciding what is still coming.

Three requirements install from git, and one of them is a UI toolkit fork

The requirements file is long, and the install method is not consistent across it.

Most entries are ordinary version constraints, some pinned exactly, such as charset-normalizer at 2.1.1, packaging at 22.0, moviepy at 1.0.3, librosa at 0.10.1, timecode at 1.3.1 and setuptools at 69.5.1, and others bounded from below, such as transformers, spacy, scikit-image, imageio, soxr, pydantic and pyannote.audio. Several of the heaviest are unpinned entirely: torch, torchaudio, torchvision, numpy, onnxruntime, tokenizers, dtw-python, resampy, darkdetect and ftfy.

Three entries are not packages at all. openai-whisper is installed from the Whisper repository at a named commit, CLIP is installed from the OpenAI CLIP repository at another named commit, and CustomTkinter is installed from a fork under the maintainer's own account at a third.

That last one has a lasting consequence. The desktop interface depends on a personal fork pinned to a single commit, so upstream fixes and upstream releases only arrive when the fork is updated, and a user installing from this file is running code that is one person away from the published toolkit.

Three things leave your machine, and one of them runs on every start

The privacy section is more specific than most, and it is worth reading in the order it is written.

The claim first: the tool does not send anything you do not want to the internet, and it uses your local machine to transcribe and translate audio.

Then the enumeration. The first item is the StoryToolkitAI API key check, sent to storytoolkit.ai, and it happens only when the key is entered in the settings window. The second is the Assistant, which talks to OpenAI, to storytoolkit.ai, or to another external provider, and only for the contexts and messages you select and send.

The third item is the one that is easy to skim past. The tool checks for updates on every start.

So the Assistant is opt-in and scoped, the key check happens once at configuration, and the update ping is automatic and unconditional. That is a normal arrangement for a desktop tool, and it is the only part of the list a reader who cares about offline use needs to act on.

Local features stay free, the Assistant needs a provider, and Patreon funds the rest

There is a section headed with the question of whether the tool is really free, and the answer is qualified rather than simple.

Yes, for the core. It runs locally and no additional account is needed to transcribe, index video or search, and those features are said to always be free as long as your machine supports them without external services. That is the honest condition: the promise is tied to the hardware rather than to a policy.

No for the Assistant when you want to run it against an external provider such as OpenAI. The workaround is the same as anywhere else in the model world, running local LLMs instead, and the compatibility list names ollama, vLLM and LM Studio alongside OpenAI.

Development is funded by Patreon, and the page is described twice, once as support and once as a way to get access to new features earlier. That last phrase is the actual product split: supporters see features before the release page does, which lines up with the note that standalone builds always lag the git version.

Editorial conclusion

StoryToolkitAI is a genuine local tool rather than a wrapper, and the privacy story is the strongest thing about it: transcription, translation and search stay on the machine, and only the Assistant needs a provider. The practical question is which build you get. The standalone release is well behind the branch, several features are marked as belonging to the non-standalone version until a future release, and a fork of the UI toolkit is pinned to one commit. Install from source, read FEATURES.md before trusting a checkbox, and decide in advance how much of your footage you are willing to send to an external model.

Frequently asked questions

how to use storytoolkitai

StoryToolkitAI installs from the standalone release on the releases page or from source, with the steps in INSTALLATION.md. Transcription, indexing and search need no account and run locally, and the feature list is versioned: the story editor from 0.20.1, transcript translation and AI story creation from 0.22.0, and automatic speaker detection from 0.23.0.

What does StoryToolkitAI do with footage?

It transcribes it locally, indexes scenes so you can search content without typing exact words, detects speakers, groups transcript lines, and uses large language models to select footage and assemble stories. Results export as EDL, XML or Fountain, plus transcript formats including SRT, TXT and AVID DS.

Does StoryToolkitAI send my footage or transcripts to the internet?

Not for the core work. Transcription, translation and search run on your own machine. Three things do leave it: the API key check to storytoolkit.ai when you enter a key, the Assistant when you select a context or message to send to OpenAI, storytoolkit.ai or another provider, and an update check that runs on every start.

What does StoryToolkitAI need installed from git?

Three requirements are not packages: OpenAI Whisper at a named commit, OpenAI CLIP at a named commit, and CustomTkinter from a fork under the maintainer's account at a named commit. That last one means the desktop interface depends on a personal fork, so upstream toolkit fixes arrive only when the fork moves.

Can StoryToolkitAI work without DaVinci Resolve?

Yes. The tool works locally and independently of any editing software, exporting to EDL, XML and Fountain instead. The Resolve integration, which needs DaVinci Resolve Studio 18 or newer, adds timeline marking and navigation through the transcript, AI search over timeline markers, copying markers both ways, and direct import of subtitles into the bin.

Official sources

  1. Issues
  2. License: GPL-3.0
  3. octimot/StoryToolkitAI on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/octimot-storytoolkitai.svg)](https://hysenlabs.com/projects/octimot-storytoolkitai)