WhisperJAV: a Japanese ASR subtitle generator built around Whisper's known failure modes
ASR/STT subtitle generator. Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD. Noise-robust for JAV
At a glance
- What is it?
- WhisperJAV is an MIT-licensed Python tool that turns Japanese adult video audio into SRT files on your own machine, using scene segmentation, VAD clamping and Japanese-specific post-processing instead of a single ASR model.
- Who is it for?
- Adopt WhisperJAV if your source is Japanese adult video, your audio is noisy or long-form, and you want the file to stay on your own disk; the pipeline is explicitly built for that mismatch and the post-processing pass is where most of the domain work sits. Do not adopt it if you need a general-purpose multilingual transcriber, if you cannot install FFmpeg and a Python 3.10 to 3.13 environment, or if you need documented rollback and upgrade procedures before you commit.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 14 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The mismatch WhisperJAV is written against
Whisper-family models are trained on clean, curated speech. The README states the opposite for this domain and names three concrete ways the assumptions break. First, the acoustic profile: low signal-to-noise ratio, a high density of non-verbal vocalisations whose spectra mimic real Japanese syllables such as fu, extreme volume swings from whispers to screams, and theatrical role language (yakuwarigo) that is absent from training corpora. Second, long-form drift: these are feature-length recordings, and over long stretches of ambiguous audio the model's attention collapses and it fills the void with repeated or invented text. The README cites this as a documented failure mode of Whisper-family models. Third, what it calls the pre-processing paradox: blanket denoising and vocal isolation can strip the high-frequency detail needed to tell consonants apart, and fine-tuning on this domain overfits because good datasets are scarce.
The intended user is someone with a local video file who wants Japanese subtitles without uploading the media anywhere. The README is direct about that: free, runs on your own machine, no cloud upload. It is not a general transcription service and it does not pretend to be one. The scope is Japanese speech in adult video, and every design decision in the pipeline follows from that scope rather than from a desire to be a universal ASR front end.
Scene detection, VAD clamping and the ChronosJAV split
The README gives the pipeline as a flowchart: audio extraction, scene detection, optional speech enhancement, speech segmentation by VAD, the ASR model, post-processing, then the .srt file. Scene detection cuts the media by predicted scenes so downstream stages receive chunks with similar acoustic character. Speech enhancement is off by default and, per the README's own pre-processing paradox, is meant to be used surgically per scene rather than as a blanket pass. VAD then finds where speech actually is inside each scene, and the README is explicit that this choice matters more than most settings because it decides what the model hears and, in the modern pipelines, where the timestamps come from.
ChronosJAV is the part worth understanding before you pick a mode. The README states that some of the strongest recognisers for this domain, including anime-whisper and Qwen3-ASR with its Japanese finetunes, do not produce reliable timestamps on their own. ChronosJAV therefore runs text generation and timing as separate stages: the VAD supplies the time skeleton and the model supplies the words. Since v1.9, timestamps come from VAD frames by default with no aligner model loaded, which the README puts at roughly 1 GB less VRAM; a Qwen forced-aligner mode remains available in the settings for word-level alignment. The same decoupling is why a new model can be added without rebuilding the pipeline, since anything that turns audio into text can be slotted in.
The post-processing pass is where the Japanese-specific work sits, and it is more opinionated than the pipeline diagram suggests. It regroups sentences with awareness of ending particles (ね, よ, わ, の), aizuchi (うん, はい) and dialect patterns such as Kansai-ben. It removes hallucination and repetition, drops subtitle lines that are purely moans or breathing kana while protecting real dialogue with an evidence check, repairs timing by pulling in the start of a line whose duration is absurdly long for its text (the end stays put, and the console reports how many lines were retimed), and resolves scene-boundary overlap. This is careful plumbing around known model weaknesses, and the README says so rather than claiming the problem is solved.
Installing WhisperJAV and a first real run
The package metadata in pyproject.toml is the single source of truth for installation. It requires Python 3.10 or newer and below 3.14, and it splits dependencies into extras so you install only what you need: cli for the command line with audio processing, gui for the graphical interface, translate for the translation module, all for everything, and colab for a Colab-optimised set. A core-only install is available but will not give you the audio processing path.
pip install whisperjav[cli]That installs the CLI extra. If you want the desktop interface instead, the gui extra is the one to pick, and the README notes a Windows installer that puts a desktop shortcut in place.
whisperjav video.mp4Running the command with just a file path uses the defaults. Subtitles land next to your video as .srt, and the README states that any input FFmpeg can read works, including MP4, MKV, AVI, WMV, MP3, WAV and FLAC. If the command fails immediately, the first thing to check is whether FFmpeg is on your PATH, since the README lists it as the input layer rather than an optional dependency.
whisperjav video.mp4 --mode balanced --sensitivity aggressiveMode and sensitivity are the two knobs you will actually touch. The README documents eight modes: balanced (Faster-Whisper, the default), fidelity (OpenAI Whisper, slowest and most thorough of the classic pipelines), fast (OpenAI Whisper plus scene detection), faster (Faster-Whisper with minimal preprocessing), qwen and anime-whisper (both ChronosJAV), transformers (HuggingFace, for Kotoba and other HF Whisper models), and crispasr (an experimental bring-your-own external build). Sensitivity applies to every mode and has three values: conservative for fewer false positives on noisy content, balanced, and aggressive for catching more quiet dialogue in whisper or ASMR material, which the README names as the tuning target of most of its benchmark work.
whisperjav /path/to/folder --output-dir ./subtitlesPassing a directory processes the whole folder, and --output-dir puts the results somewhere other than next to the source. Output defaults to SRT; --output-format both writes WebVTT as well. The GUI path is the same pipeline: launch WhisperJAV from the desktop shortcut, or run whisperjav-gui, then add files, pick a mode and click Start.
Where WhisperJAV is the wrong tool
The README states plainly that results still vary with source audio quality, and that sentence deserves more weight than the mode table around it. There are no accuracy figures in the documentation, no per-mode benchmark numbers, and no statement about which mode wins on which kind of file. Mode selection is therefore an empirical exercise on your own audio, not something you can read off a table. If you need a documented accuracy target before you commit, this project does not give you one.
The second limitation is scope. Everything domain-specific here is Japanese-specific: the sentence regrouping keys on ending particles, aizuchi and dialect patterns, and the sound-only line removal is tuned to drop moans and breathing kana. Point this at English or Spanish audio and the post-processing pass has nothing useful to do, while the scene detection and VAD stages still cost you time. For general multilingual transcription, a plain Whisper or Faster-Whisper setup is the simpler choice, and the README's own mode list shows WhisperJAV wrapping those engines rather than replacing them.
The third is operational. The package metadata classifies the project as Development Status 4 - Beta. The crispasr mode is labelled experimental and requires you to bring your own CrispASR build. The README does not document rollback, does not describe an upgrade path between releases, and does not state what happens to your configuration when you move from one version to the next. If your workflow needs a supported upgrade procedure, that is a gap you have to accept or work around.
How this differs from running Whisper directly
The obvious alternative is Faster-Whisper or OpenAI Whisper on their own, and the difference is not the model. WhisperJAV's balanced mode uses Faster-Whisper and its fidelity mode uses OpenAI Whisper, so at the recognition step you are running the same engines. The difference is everything around them: scene-based segmentation so the model never processes a mixed acoustic stream, VAD clamping with measured padding as the main defence against hallucination on non-speech, tuned confidence thresholds that discard low-quality output, and a Japanese-aware cleanup pass.
That framing matters for the decision. If your audio is clean, short and clearly spoken, direct Whisper usage gives you the same text with less machinery and fewer settings to get wrong. WhisperJAV earns its complexity specifically on long, noisy recordings where attention collapse and repetition are the failure you are fighting. The README's own account of the pre-processing paradox is a good example of the difference in posture: a naive pipeline denoises first because it seems obviously right, and this one keeps enhancement off by default and applies it per scene, because blanket denoising can remove the detail the recogniser needs.
A second point of comparison is ChronosJAV against the classic pipelines. Classic Whisper modes produce timestamps as part of recognition. ChronosJAV separates the two, taking the time skeleton from VAD and the words from the model, which is what makes Qwen3-ASR and anime-whisper usable here at all. The trade-off is that timing quality now depends on the VAD rather than on the recogniser, and the README notes that the default since v1.9 loads no aligner model, saving roughly 1 GB of VRAM at the cost of word-level alignment unless you switch the Qwen forced-aligner mode back on in the settings.
Licence and what maintenance costs you
WhisperJAV is MIT-licensed, both in the repository metadata and in pyproject.toml. That is permissive: you can use, modify and redistribute it, including in commercial contexts, provided the licence notice is preserved. The repository also carries a NOTICES file alongside LICENSE, which suggests third-party components are tracked separately, and those components carry their own terms. The ASR models the modes pull in, including Qwen3-ASR, anime-whisper and the HuggingFace Whisper variants, are downloaded and licensed independently of WhisperJAV itself, so the MIT licence on this code does not settle the terms of the models you run through it. Check the model licences for your own use case; this is not legal advice.
On maintenance, the last push to the default branch was on 2026-09-05, and the most recent release in the repository is v1.9.0 on 2026-08-15, which the release notes describe as adding FireRedVAD, a QwenASR finetune, CrispASR and improved subtitle timing. Before that, v1.8.14 arrived on 2026-05-10 and v1.8.13 on 2026-05-05. The repository is not archived. The upgrade cost is concentrated in the dependency extras: because pyproject.toml is the single source of truth and splits dependencies into cli, gui, translate, all and colab, a version bump can change what an extra pulls in without changing the command you type. The README does not describe a migration guide between releases, so pinning a version you have validated is the conservative move if you depend on a specific mode's output.
Editorial conclusion
Adopt WhisperJAV if your source is Japanese adult video, your audio is noisy or long-form, and you want the file to stay on your own disk; the pipeline is explicitly built for that mismatch and the post-processing pass is where most of the domain work sits. Do not adopt it if you need a general-purpose multilingual transcriber, if you cannot install FFmpeg and a Python 3.10 to 3.13 environment, or if you need documented rollback and upgrade procedures before you commit. Verify first that a mode actually fits your audio: run the same short clip through --mode balanced and --mode qwen with --sensitivity aggressive and compare the SRT, because the README states results vary with source audio quality and the mode table gives no accuracy numbers.
Frequently asked questions
How do I install WhisperJAV?
Install it with pip using one of the extras defined in pyproject.toml, for example pip install whisperjav[cli] for the command line with audio processing or whisperjav[gui] for the graphical interface. It requires Python 3.10 or newer and below 3.14. A Windows installer is also mentioned in the README, which adds a desktop shortcut.
Which WhisperJAV mode should I use for noisy audio?
The README describes conservative sensitivity as producing fewer false positives and being suited to noisy content, and balanced as the default mode with a good speed and accuracy balance. The README also states that results vary with source audio quality and gives no per-mode accuracy numbers, so the choice is best settled by comparing output on your own clip.
Does WhisperJAV upload my video anywhere?
No. The README states that it runs on your own machine with no cloud upload of your media. Colab and Kaggle notebooks are offered as alternative environments, but the local install path keeps the file on your disk.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/meizhong986-whisperjav)