Model or dataset
byjlw/video-analyzer avatar
byjlw/video-analyzer

video-analyzer describes a video by describing every key frame in order

Analyze videos using LLMs, Computer Vision and Automatic Speech Recognition

1,592 stars221 forksPythonApache-2.0

At a glance

What is it?
A Python command line tool that samples frames with OpenCV, sends each one to a vision model together with the previous frames, and merges the chain with a Whisper transcript. It runs against Ollama locally or any OpenAI-compatible endpoint, and the two version numbers in its own files do not agree.
Who is it for?
video-analyzer is a small CLI with real plumbing behind it, and the two things to settle before trusting it are both version questions. One is Python: the guide asks for 3.11, the packaging file allows 3.8, and the dependency list carries a marker for 3.13.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 166 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The guide asks for Python 3.11 and the packaging file allows 3.8

Two files, two floors. The requirements section opens with Python 3.11 or higher. The packaging file says:

python
    entry_points={
        "console_scripts": [
            "video-analyzer=video_analyzer.cli:main",
        ],
    },
    python_requires=">=3.8",
    include_package_data=True

Nothing in between reconciles them. pip reads the lower number and will install on 3.8, 3.9 and 3.10 without complaint, and on those versions you get a tool whose own documentation says it is unsupported. The entry point in the same block is worth noting too: the installed command is video-analyzer, wired straight to video_analyzer.cli:main, so there is no separate launcher to configure.

The dependency list adds a third data point rather than settling the question. Its last few entries include audioop-lts behind a marker for Python 3.13, which is a backport for a module that left the standard library, and both whisper packages are pinned side by side. A tool that supports 3.8 through 3.13 in one dependency file and claims 3.11 in its requirements section is telling you the packaging metadata has drifted from the documentation.

Two Whisper implementations are declared at once

The dependency list is short and has one entry that explains why the project needs a compatibility note:

code
opencv-python>=4.8.0
numpy>=1.24.0
torch>=2.0.0
openai-whisper>=20231117
requests>=2.31.0
Pillow>=10.0.0
pydub>=0.25.1
audioop-lts>=0.2.2; python_version >= "3.13"
faster-whisper>=0.6.0

openai-whisper and faster-whisper are both unconditional. They are two implementations of the same job, and neither is marked optional, so a plain install of the package pulls in both. The transcription stage is the one place where the tool claims to handle bad audio, and the design notes say it does that with confidence checks on the Whisper output rather than by skipping segments, so the behaviour of that stage depends on which of the two is doing the work.

pydub sits between them and the audio file, and pydub is what reaches for the audioop module. That is why audioop-lts is there at all, and why its marker starts at 3.13. Everything else in the list is ordinary: OpenCV for frame extraction, torch as the runtime underneath whatever you run locally, Pillow for images, requests for the API client.

The output path is written with a backslash

The output section says the tool generates a JSON file at output\analysis.json. The backslash is a Windows path separator, and it appears in the same document whose install section gives separate commands for Ubuntu and Debian, macOS and Windows.

On Linux and macOS a backslash is an ordinary character in a filename, so taken literally that path is one file called output\analysis.json in the current directory rather than a file called analysis.json inside an output directory. This is a documentation slip rather than a broken code path, and the rest of the output description is consistent with the file existing: metadata about the analysis, the audio transcript if available, the frame-by-frame analysis, and a final video description.

The sample output that follows does not match that description either. Where the format is JSON, the sample is a paragraph of prose about a person in a pink t-shirt standing in front of a black plastic tub, with literal \n\n sequences and a run of dots standing in for the rest. The complete example is said to be in docs/sample_analysis.json, which is the file to read if you want to know what the JSON actually looks like.

Every frame analysis is handed the frames before it

The pipeline has three stages. Frame extraction and audio processing use OpenCV to pull key frames and Whisper to transcribe, with confidence checks on poor quality audio. Frame analysis then sends each frame to the vision model. Video reconstruction combines the frame analyses in order, folds in the transcript, and uses the first frame to set the scene.

The middle stage is the one with a design decision in it. Each analysis includes context from previous frames, and the sequence is kept chronological, which is how a set of independent captions becomes a description of a scene rather than a list of unrelated snapshots. It also means the frames are not independent: the second call carries the first call's output, the third carries both, and the context grows along the chain.

That has consequences the guide does not spell out. The number of vision model calls scales with the number of key frames, and each one is more expensive than a single-frame request. Nothing in the requirements section names a default sampling interval or a maximum frame count, so the knob that controls both cost and context growth is documented in docs/USAGES.md rather than here. The prompt template driving that stage is a file, frame_analysis.txt, shipped as package data alongside the packaged config JSON.

The same 11B vision model runs on your hardware or on OpenRouter

Two routes, one model. Locally you install Ollama, pull the default vision model with ollama pull llama3.2-vision, and start the service with ollama serve. The hardware requirement for that route is stated plainly: at least 16GB of RAM with 32GB recommended, and either a GPU with at least 12GB of VRAM or an Apple M Series machine with at least 32GB.

The other route keeps the model and changes the transport. Any OpenAI-compatible service is selected with three flags, as in the cloud example from the quick start:

bash
video-analyzer video.mp4 \
    --client openai_api \
    --api-key your-key \
    --api-url https://openrouter.ai/api/v1 \
    --model meta-llama/llama-3.2-11b-vision-instruct:free

The same values can sit in config/config.json instead of on the command line, under a clients object with a default name and a per-client block holding api_key and api_url.

The feature list opens by saying it can run completely locally with no cloud services or API keys needed. Both claims are true and they are the same model: llama 3.2 11b vision is what Ollama serves, and the :free suffix in that command is the same model reached through someone else's endpoint. So the choice is not local model versus cloud model. It is whose hardware and whose bill.

Prompt tuning is a second package released from the same repository

The optimization path is not in the main package. It lives in its own top-level directory, installs on its own line:

bash
pip install video-analyzer-tune

and it works by having you run video-analyzer on a few representative videos, edit the outputs to show what good results look like, and then letting DSPy MIPROv2 find better prompt instructions. The tuned prompts are written out as new files that you point to from your config, and the guide states plainly that the main package is unaffected.

That separation is the right call, because the tuner pulls in a different stack from the analyzer. It also means the repository releases two things under two tag schemes: a v0.1.2 tag for the analyzer and a tune-v0.1.0 tag for the tuner, both dated the same day. The analyzer's version lives in its packaging file, which declares 0.1.2, so a version check means reading two different artefacts depending on which of the two you installed.

A UI directory sits in the tree and never appears in the guide

The top-level entries include video-analyzer-ui/ and video-analyzer-tune/. Only the second is discussed anywhere in the documentation, in a section of its own. The UI directory has no mention: not in the table of contents, not in the project structure, not in the install steps, which only ever produce a console script.

There is also test_prompt_loading.py sitting at the root of the repository rather than in a tests directory, and no test dependency appears in requirements.txt, so the one test file in the tree has no declared way to be run.

Both facts point at the same thing: the tree and the documentation are maintained separately, and the last recorded change, dated 2026-04-19, came a month after the 0.1.2 tag. So what is on the default branch is ahead of what the guide describes. Before assuming the CLI is the whole product, or that the guide matches the code, look at the tree.

Two Design headings, one of which points at a file

The document has a Design section near the top, where the three stages are described, and a second Design section further down that exists only to send you to docs/DESIGN.md for implementation detail and for how to make changes. The table of contents lists Design once. Two headings with the same name in one file is a small thing, but it means the anchor for the design overview and the anchor for the design link are different targets, and any link written against one will land on the other.

The rest of the documentation splits the same way. The usage guide, the design detail, the contributing guidelines and the complete sample output are all files under docs/, and the README is largely an index to them. Configuration is delegated the same way: three layers in priority order, command line arguments first, then a user config at config/config.json, then the packaged defaults.

That last one has a wrinkle worth knowing. The user config lives at config/config.json relative to where you run the command, while the packaging file ships config/*.json as package data inside the video_analyzer package. Two files with the same relative name, one of which is yours and one of which is not, and nothing in the guide says which one wins when both exist.

Editorial conclusion

video-analyzer is a small CLI with real plumbing behind it, and the two things to settle before trusting it are both version questions. One is Python: the guide asks for 3.11, the packaging file allows 3.8, and the dependency list carries a marker for 3.13. The other is where your frames and audio go. Running locally means an 11B vision model on your own hardware; pointing the same model at OpenRouter changes only the transport, not the model, and the frames and transcript go to someone else's endpoint. Read docs/USAGES.md for the sampling options the guide never states, diff the Python floor before you install, and treat video-analyzer-tune as separate software with its own tag series.

Frequently asked questions

What does video-analyzer need in order to run locally?

Python 3.11 or higher per the requirements section, FFmpeg for audio processing, and Ollama with llama3.2-vision pulled. For local models the guide asks for at least 16GB of RAM, 32GB recommended, plus a GPU with at least 12GB of VRAM or an Apple M Series machine with at least 32GB.

Can video-analyzer use a cloud model instead of Ollama?

Yes, through any OpenAI-compatible service, selected with --client openai_api, --api-key, --api-url and --model, or set under a clients object in config/config.json. The guide notes that llama 3.2 11b vision can be used on OpenRouter for free by adding :free to the model name.

Where does video-analyzer write its results?

A JSON file containing analysis metadata, the audio transcript if available, the frame-by-frame analysis and a final video description. The guide writes the path as output\analysis.json, using a backslash separator.

Which Python versions does video-analyzer support?

The two numbers disagree. The requirements section asks for Python 3.11 or higher, while the packaging file declares python_requires of 3.8 or higher, so pip will install on versions the guide does not list. The dependency list also carries a Python 3.13 marker for a backported audio module.

How does video-analyzer choose and describe key frames?

OpenCV extracts the key frames and each one is analysed by the vision model, with every analysis including context from the previous frames so the sequence stays chronological. A third stage combines the frame analyses with the audio transcript and uses the first frame to set the scene.

Can the prompts used by video-analyzer be tuned?

Yes, through a separate package installed with pip install video-analyzer-tune, which uses DSPy MIPROv2 to improve the frame analysis and reconstruction prompts. Tuned prompts are written as new files that you point to from your config, so the main package is left alone.

Official sources

  1. byjlw/video-analyzer on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/byjlw-video-analyzer.svg)](https://hysenlabs.com/projects/byjlw-video-analyzer)