ffsubsync: aligning subtitles to video with speech detection and FFT
Automagically synchronize subtitles with video.
At a glance
- What is it?
- ffsubsync is a Python command-line tool that aligns a subtitle file to a video or to a correctly synced reference subtitle. It is language-agnostic, ships an MIT licence, and the README also points to a browser version that needs no install.
- Who is it for?
- ffsubsync suits people who already have a subtitle file and a video whose timing does not match, and who are comfortable with a command line or Docker. It is the wrong tool when no subtitle exists at all, when the subtitle is in a different language and no correctly synced reference is available, or when the desync is not a timing problem.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 69 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem ffsubsync targets: subtitles that arrive at the wrong moment
A subtitle file can be perfectly translated and still be useless, because it starts a few seconds before or after the dialogue it describes. The README frames the whole project around this single defect: aligning subtitles to the correct starting point within the video. That framing matters, because it tells you what the tool does not attempt. It does not translate, retime individual lines to match a speaker, or repair a subtitle whose text is wrong. It finds a shift, and where relevant a framerate ratio, between two timelines.
The audience is the person holding a video and a mismatched .srt. The README's animated example uses the same clip twice, once with subtitles arriving early and once corrected, which is a fair picture of the intended use. The project describes itself as language-agnostic, and that is the design decision that shapes everything else: instead of parsing the words, it compares the pattern of speech in the audio against the pattern of subtitle events. Text in any script can be aligned this way, provided the subtitle timings roughly correspond to speech.
The repository is not archived, and the last push was on 2026-07-24. That is recent enough that the project is not abandoned, but the setup.py classifiers still carry Development Status 3 - Alpha, which is worth reading literally: the interface has grown (0.5.0 and 0.5.1 both landed in mid-2026) and flags have been added over time.
How the alignment works: voice activity detection against subtitle timings
The mechanism is a comparison of two binary-ish signals. On one side, ffsubsync extracts speech from the reference. On the other, it derives speech from the subtitle file. It then searches for the offset, and if needed a scaling factor, that best lines the two up.
The reference decides which branch runs. The README states that ffsubsync uses the file extension to decide whether to perform voice activity detection on the audio or to directly extract speech from an srt file. So a .srt reference goes down the subtitle path, and a video or audio file goes down the audio path, where ffmpeg decodes the stream and a voice activity detector marks where speech occurs. The topics on the repository list voice-activity-detection, fast-fourier-transform and string-alignment, which is consistent with a pipeline that turns both sides into sequences and then aligns them.
The default VAD is webrtcvad-wheels, pinned in requirements.txt. setup.py declares an optional extra, torch, described as the optional dependency for the silero / fused VAD backends, so a second and third detector exist for people willing to install PyTorch. The README does not explain when to prefer them, which is a documentation gap rather than a design flaw.
Because the search covers a framerate ratio as well as a constant offset, a subtitle written for 23.976 fps and played against 25 fps material is a case the tool is built to handle. The README makes this explicit for --multi-segment-sync: a framerate mismatch is still detected and corrected.
Installing ffsubsync and running a first sync
The README puts ffmpeg first, before the Python package, and that ordering is not decorative: audio references are decoded by ffmpeg, and without it on your path the audio branch has nothing to work with. On macOS the documented command is brew install ffmpeg. Windows users are told to make sure ffmpeg is on the path and can be referenced from the command line. The package itself is compatible with Python >= 3.6 according to the README, while the classifiers in setup.py list 3.6 through 3.10.
brew install ffmpeg
pip install ffsubsyncThe README notes a second install path, pip install git+https://github.com/smacke/ffsubsync@latest, which it introduces with the phrase "if you want to live dangerously". Treat that as a warning about tracking the branch rather than a release.
Three entrypoints are registered: ffs, subsync and ffsubsync. They are the same program. The canonical invocation names the reference first, then the unsynchronized subtitle, then the output.
ffs video.mp4 -i unsynchronized.srt -o synchronized.srtIf you omit -i, the tool looks for subtitles sitting next to the reference that share its name, so ffs video.mp4 picks up files like video.srt and video.en.srt and writes video.synced.srt for each, leaving the originals untouched. The README states that previously-synced *.synced.srt outputs are skipped, so re-running the same command is safe, and that --overwrite-input replaces the detected files in place instead. Auto-detection is skipped when subtitles arrive on stdin.
The reference does not have to be local. Anything ffmpeg can read works, including http(s)://, rtmp://, rtsp:// and ftp:// URLs, and a remote .srt can serve as the reference too. The README is candid about the cost: processing streams the reference over the network, so it depends on connection stability, and for large or flaky sources downloading first is more reliable. Two flags address that. --max-duration-seconds N processes only the first N seconds measured from --start-seconds, and because ffmpeg stops reading once that duration is reached, it also stops downloading. --extract-audio-first copies the remote audio track to a local temp file without re-encoding, which the README recommends on flaky connections; it is ignored for local references.
The reference subtitle path, and why it is the most reliable mode
When you already have a subtitle in a language you do not read, and an unsynchronized subtitle in one you do, the README's advice is to use the synced file as the reference instead of the video. That is the strongest configuration the tool offers, because it removes audio decoding, VAD and the acoustic noise floor from the problem entirely. Both sides become subtitle event sequences, and the alignment is a pure timing comparison.
ffsubsync reference.srt -i unsynchronized.srt -o synchronized.srtThe trade-off is trust. The reference is assumed correct, and the README does not describe any validation of that assumption. If your reference is itself shifted, the output inherits the shift. There is also a harder limitation that follows from how the tool works: it aligns speech patterns, not text. Two subtitle files for the same video in different languages will not have identical cue boundaries, since a translation changes line lengths and split points, so the match is approximate. The README does not discuss accuracy expectations for cross-language subtitle references, and that silence is worth noting before you build a workflow around it.
Long references: --multi-segment-sync and its costs
For a two-hour film, --max-duration-seconds is a blunt instrument, because it only looks at the beginning. Desync that appears later, or a drift that accumulates, is invisible to it. --multi-segment-sync exists for that case: it samples several short segments spread across the whole reference and runs speech detection on just those.
The README is specific about the consequences. Only the sampled audio is extracted, and for remote references only that audio is downloaded. Because each segment keeps its true position on the timeline, the framerate-ratio and offset search is unchanged, so a framerate mismatch is still detected and corrected. Three flags tune it: --segment-count N with a default of 8, --skip-intro-outro which skips the first 30 seconds and the last 60 seconds on the grounds that they often lack dialogue, and --parallel-workers N which overlaps segment downloads with a default of 4. It applies to video and audio references only, so it does nothing in the subtitle-reference mode described above.
ffs "https://example.com/video.mp4" -i unsynchronized.srt -o synchronized.srt --multi-segment-syncThe honest limitation is that sampling is sampling. A segment that lands in a stretch of music, silence or crowd noise contributes little, and the README's own rationale for --skip-intro-outro shows the authors are aware that not every part of a file carries usable speech.
Docker as the escape hatch from Python and ffmpeg setup
The Dockerfile builds in two stages from python:3.14-slim, installs ffmpeg with apt-get, creates a virtualenv at /opt/venv, installs the package, runs ffs --version as a build-time check, and sets WORKDIR to /video with ENTRYPOINT [ "ffs" ]. That layout is the whole usage contract: mount the directory containing your files at /video, and the arguments you pass become ffs arguments.
docker pull ghcr.io/smacke/ffsubsync:latest
docker run --rm -v "$PWD":/video ghcr.io/smacke/ffsubsync:latest \
video.mp4 -i unsynchronized.srt -o synchronized.srtThis is the cleanest answer for Windows users who do not want to fight ffmpeg on the path, and for anyone who does not want a Python environment on the host. It is also the slowest path for a one-off file, since the image carries Python, ffmpeg and the dependency set. Note that the image installs a released version by default; the Dockerfile accepts an FFSUBSYNC_VERSION build argument, which is a build-time choice, not a runtime flag.
Where ffsubsync is the wrong tool, and how it compares to alass
Three cases defeat it. First, no subtitle at all: the tool aligns an existing file, and the README never suggests it can generate one from audio. Second, a subtitle whose cue timings bear no relation to speech, for example one built from a different cut of the film with scenes reordered, where no single offset or ratio explains the difference. Third, a reference you cannot trust, since the tool takes the reference as ground truth.
There is a fourth, quieter case: the README does not document rollback or a dry-run mode. The safe-by-default behaviour of writing <name>.synced.srt and skipping existing outputs is a reasonable substitute, but if you pass --overwrite-input you are overwriting the detected files in place, and the README says nothing about a backup.
The obvious alternative is alass, which people search for alongside this project. The approaches differ at the point of input: ffsubsync can derive its timing signal from the audio through voice activity detection, which means it works even when you have no correctly synced subtitle in any language. alass aligns subtitle against subtitle, so it needs a reference subtitle and never touches the audio track. If you have a good reference in a language you cannot read, both tools are in play and the difference is mostly implementation. If you have only the video, ffsubsync is the one with a path forward. The project also publishes a browser version at smacke.github.io/ffsubsync, which the README says syncs against either a correctly synced reference subtitle or a video or audio file, decoding audio in-browser with ffmpeg.wasm and reading large files lazily without uploading them. That is a genuine option for a single file on a machine where you cannot install anything.
Licence, dependencies and the cost of keeping it current
The licence is MIT, declared in setup.py and shown in the README badge. MIT is permissive: it allows commercial and closed-source use, modification and redistribution, provided the copyright notice and permission notice are kept. It is also a licence with no patent grant and no copyleft obligation. That is a summary of the licence text, not legal advice; if the tool ends up inside a product you ship, have someone qualified read the actual LICENSE file.
The runtime dependency list in requirements.txt is broader than the single ffmpeg requirement suggests: auditok, chardet, charset_normalizer, faust-cchardet, ffmpeg-python, numpy, pysubs2, rich, srt, tqdm, typing_extensions and webrtcvad-wheels, with pysubs2 version-gated on Python below 3.7. Several of these are pinned or lower-bounded rather than fully pinned, so a fresh install can pull newer versions than the ones the project was last tested against. The optional torch extra is not installed by default and is only needed for the silero and fused VAD backends.
Upgrade cost is low in normal use: pip install --upgrade ffsubsync, or pull ghcr.io/smacke/ffsubsync:latest again. The real maintenance question is not code churn but environment churn, since ffmpeg is installed separately and its behaviour is outside the project's control. The last push was on 2026-07-24, and the project is not archived.
Editorial conclusion
ffsubsync suits people who already have a subtitle file and a video whose timing does not match, and who are comfortable with a command line or Docker. It is the wrong tool when no subtitle exists at all, when the subtitle is in a different language and no correctly synced reference is available, or when the desync is not a timing problem. Before adopting it, verify that ffmpeg is on your path, that your Python is 3.6 or newer, and that the reference you intend to use is genuinely in sync, because the tool trusts it without checking.
Frequently asked questions
How do I install ffsubsync?
Install ffmpeg first, for example with brew install ffmpeg on macOS, then run pip install ffsubsync. The README states the package is compatible with Python 3.6 and newer. A Docker image is also published at ghcr.io/smacke/ffsubsync:latest if you would rather not install Python locally.
How do I use ffsubsync to sync an srt file with a video?
Run ffs video.mp4 -i unsynchronized.srt -o synchronized.srt, naming the reference first and the unsynchronized subtitle with -i. If you omit -i, ffsubsync looks for subtitles next to the reference that share its name and writes a .synced.srt file for each, leaving the originals untouched.
What is the difference between subsync and ffsubsync?
The README registers ffs, subsync and ffsubsync as three entrypoints for the same program, so on the command line there is no difference between them. The project name is ffsubsync; subsync is just a shorter alias.
Can ffsubsync sync subtitles automatically?
Yes, within limits. It compares speech detected in the reference against the timing of subtitle events and searches for the offset, and where needed a framerate ratio, that aligns them. The reference can be a video or audio file, or an already correctly synced subtitle file.
How do I fix subtitle sync in VLC?
ffsubsync does not control VLC; it produces a corrected subtitle file that you then load in the player. The README's workflow is to run ffs against the video or a synced reference subtitle and write a synchronized .srt, which replaces the mistimed file you were loading.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/smacke-ffsubsync)