FunClip: FunASR-based video transcription and subtitle clipping with a local Gradio UI
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
At a glance
- What is it?
- FunClip is an MIT-licensed Python tool that transcribes video with FunASR models, lets you pick text segments or speakers, and exports both clips and SRT subtitles from a local Gradio interface. Its LLM clipping path and MOSS diarization mode are the interesting parts; the model downloads and the FunASR version floor are the parts that will bite you first.
- Who is it for?
- Adopt FunClip if you are editing Chinese-language video and want transcription, speaker-aware segment selection and SRT export inside one local Gradio app, with an MIT licence and no per-minute API cost. Do not adopt it if you need a hosted service, if you cannot run a Python environment with FunASR model weights, or if your primary language is English and you want a mature English-first workflow rather than the '-l en' flag.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 14 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What FunClip actually solves for video editors
Cutting a clip out of a long recording usually means scrubbing a timeline by ear. FunClip replaces that with text. It runs speech recognition over the video with FunASR's Paraformer series models, returns the transcript with timestamps, and then lets you select text segments (or a speaker) and press a button to get the corresponding video cut. The README describes it as a "fully open-source, locally deployed automated video clipping tool", and that local deployment is the point: nothing is uploaded to a service, and there is no per-minute transcription bill.
The audience is narrow but real. If you produce Chinese-language video, interviews, lectures, or podcast footage and you already work from a Python environment, the combination of Chinese ASR quality plus speaker labelling plus SRT output covers a workflow that otherwise needs three separate tools. The homepage points at a Hugging Face Space for a quick look before installing anything, which is the honest way to evaluate whether the recognition quality is good enough for your audio.
How FunASR, CAM++ and the LLM path fit together
The pipeline has three recognisable stages. First, ASR: the default path uses Paraformer-Large, which the README says predicts timestamps in an integrated manner rather than running a separate forced-alignment step. Second, speaker attribution: FunClip integrates the CAM++ speaker recognition model, so the transcript can carry a speaker ID and you can clip everything said by one person. Third, selection and export: you choose segments in the Gradio UI, and FunClip returns both the clipped video and SRT subtitles, for the full video and for the selected segment.
Two optional paths change that shape. SeACo-Paraformer adds hotword customisation, so you can supply entity names before recognition to improve their accuracy. The v2.2.0 release added a third-party MOSS-Transcribe-Diarize path for long-form ASR that produces timestamps and anonymous speaker labels without external VAD or speaker models, which matters because it removes two moving parts from the stack. Separately, the README's highlights mention LLM-based clipping as an active direction, and the requirements list includes openai, g4f and dashscope, so the LLM path routes through one of those providers rather than a bundled local model.
Installing FunClip and running your first clip
The basic install is a clone plus a requirements install. The README gives exactly this sequence, and it is the whole Python-side setup:
# clone funclip repo
git clone https://github.com/modelscope/FunClip.git
cd FunClip
# install Python requirments
pip install -r ./requirements.txtIf you want a pinned snapshot rather than the default branch, the v2.2.1 release publishes FunClip-2.2.1.tar.gz and FunClip-2.2.1.zip alongside a SHA256SUMS file. The README is explicit that model weights are downloaded separately when FunClip starts and are not inside those archives, so a first launch on a slow connection will spend its time fetching weights.
One version constraint deserves attention before you start. The README states that the current model and subtitle compatibility paths require funasr>=1.4.9, and that installations predating this requirement should run the upgrade explicitly:
pip install -U "funasr>=1.4.9"The requirements file also pins moviepy==1.0.3, numpy==1.26.4 and gradio>=4.31.3,<5.0, so an existing environment with a newer numpy or Gradio 5 will conflict. Start in a fresh virtual environment rather than upgrading in place.
With dependencies resolved, start the local service:
python funclip/launch.pyThe README documents several flags on that command: '-m fun-asr-nano' for the flagship Fun-ASR-Nano model, '-m sensevoice' for the multilingual SenseVoice model with emotion and audio event detection, '--model moss' for the OpenMOSS long-form path, '-l en' for English audio, '-p xxx' for the port, and '-s True' to expose the service publicly. Once it is running, open the Gradio page in a browser, upload a video, wait for the transcript, select the text you want, and clip. You should end up with a video file plus SRT output for both the full video and the selection.
Where FunClip gets in the way
The dependency surface is the first real cost. FunASR, ModelScope, torch, torchaudio, transformers, librosa and moviepy all land in one environment, and the README's own version notes show how easily that drifts: the funasr>=1.4.9 floor exists because older installs break the MOSS adapter, the speaker segment normalisation and the SenseVoice fixes. This is not a tool you install once and forget.
Second, the README is candid that English support arrived as a flag rather than a redesign. The '-l en' option exists, but the highlighted model, Paraformer-Large, is described as one of the best open-source Chinese ASR models, and the English path is documented in a single line. If your footage is English, you are using the secondary path.
Third, the roadmap has open items that map to common editing needs. The On Going list still shows unchecked entries for reverse period selection while clipping and for removing silence periods. If your workflow depends on automatically dropping dead air, FunClip does not do it yet. And because everything runs locally, throughput is bounded by your own GPU or CPU, not by a service queue.
FunClip compared with Whisper-based clippers
The obvious alternative is a Whisper-based clipping tool, and the README addresses the comparison directly in its roadmap note: it says ASR using Whisper with timestamps requires massive GPU memory, and that FunClip supports timestamp prediction for vanilla Paraformer in FunASR to achieve the same result. That is the architectural difference. FunClip gets word and sentence timing from the recognition model itself, while a Whisper-based pipeline typically needs a separate alignment stage to place subtitle boundaries accurately.
The second difference is speaker handling. FunClip ships CAM++ integration and, since v2.2.0, an alternative MOSS path that emits anonymous speaker labels without external VAD or speaker models. A generic Whisper clipper leaves diarization to you. If your source material is multi-speaker, that is the deciding factor; if it is a single narrator in English, the Whisper route is likely less friction because you skip the FunASR and ModelScope stack entirely.
Licence, maintenance and what an upgrade costs
FunClip is MIT licensed, which permits commercial and private use with the usual attribution and warranty disclaimers. That covers the FunClip code. It does not automatically cover the model weights, which are downloaded separately at runtime from ModelScope, nor the third-party services the LLM path can reach (openai, g4f, dashscope, twelvelabs). Check those terms separately; the repository's LICENSE file governs only what is in the repository.
On upkeep, the last push was on 2026-09-01, the same day v2.2.1 was tagged, following v2.2.0 on 2026-08-30 and v2.1.1 on 2026-08-03. Releases are frequent and the changelog entries are specific rather than cosmetic. The upgrade instructions are correspondingly specific: v2.2.1 tells existing installations to run 'pip install -U -r requirements.txt' before restarting, and the funasr floor tells older installs to upgrade that package explicitly. Budget for a dependency reconciliation each time you move versions, not just a git pull.
Editorial conclusion
Adopt FunClip if you are editing Chinese-language video and want transcription, speaker-aware segment selection and SRT export inside one local Gradio app, with an MIT licence and no per-minute API cost. Do not adopt it if you need a hosted service, if you cannot run a Python environment with FunASR model weights, or if your primary language is English and you want a mature English-first workflow rather than the '-l en' flag. Before committing, verify three things: that your environment resolves funasr>=1.4.9 and gradio>=4.31.3,<5.0 together; that the model weights download successfully on first launch; and that the '-m moss' or '-m fun-asr-nano' path you intend to use actually produces the timestamp and speaker output you need on a sample of your own footage.
Frequently asked questions
Is FunClip free to use?
FunClip is MIT licensed and runs locally, so there is no per-use charge from the project itself. Model weights are downloaded separately at runtime, and the optional LLM clipping path can call external providers such as openai or dashscope, which have their own terms and costs.
How do I install FunClip?
Clone the repository, change into the directory, and run 'pip install -r ./requirements.txt'. The README also states that the current compatibility paths require funasr>=1.4.9, so older installations should run 'pip install -U "funasr>=1.4.9"' before starting the service.
Does FunClip work for English audio?
Yes, the launch command documents a '-l en' flag for English audio recognition. The README's highlighted model, Paraformer-Large, is described as one of the best open-source Chinese ASR models, so English is a supported option rather than the primary path.
Can FunClip clip segments from one specific speaker?
Yes. FunClip integrates the CAM++ speaker recognition model, and the README states that users can use the auto-recognized speaker ID as the target for trimming. The v2.2.0 release also added a MOSS path that produces anonymous speaker labels without external VAD or speaker models.
Does FunClip export SRT subtitles?
Yes. The README states that FunClip supports multi-segment free clipping and automatically returns full video SRT subtitles and target segment SRT subtitles. In v2.2.1 the built-in subtitle renderer uses Pillow and a bundled font, so standard subtitle clipping no longer requires ImageMagick.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/modelscope-funclip)