unblob: recursive firmware extraction with 78+ format handlers
Extract files from any kind of container formats. unblob Accurate, fast, and easy-to-use extraction suite for binary blobs. unblob parses unknown binary blobs for 78+ archive, compression, and file-system formats**, extracts their content recursively, and carves out unknown chunks.
At a glance
- What is it?
- unblob is a Python extraction suite for unknown binary blobs, built around precise chunk boundaries and a plugin system. It is strongest on firmware images and weakest when you need a format nobody has written a handler for.
- Who is it for?
- Adopt unblob if you routinely open firmware images and want chunk offsets, entropy figures and a JSON report rather than a directory of files with no provenance. Skip it if your target format has no handler and you cannot write one, or if you need extraction to run without the external tools that the pip install leaves you to install yourself.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What unblob is for, and who actually needs it
A firmware image is usually not one archive. It is a header, then a SquashFS or JFFS2 file system, then a compressed kernel, then padding, and somewhere in the middle a vendor blob that matches nothing in any signature database. The usual workflow is to run binwalk, look at the offsets it prints, and then carve things out by hand with dd. That works until the tool reports a start offset but no end, or reports a false positive inside a compressed stream.
unblob targets that workflow specifically. The README describes it as a companion for "extracting, analyzing, and reverse engineering firmware images", and the design choices follow from that: it reports both start and end offsets for each chunk, it recurses into what it extracts, and it carves out the regions that match no known format instead of silently dropping them. The audience is security researchers, embedded developers and anyone doing incident response on a device they cannot boot.
It is not a general-purpose archive tool. If you have a zip file and want the files, unzip is faster and has no dependency list. unblob earns its complexity when the input is unknown and you need a record of what was found and where.
Chunk detection, recursion and the unknown-chunk path
The pipeline starts with format handlers. Each handler knows how to recognise a format and how to compute where it ends, which is why the README emphasises "identifies both start and end offsets of each chunk according to the format standard". A handler that can only find a start offset produces a different kind of result than one that can bound the chunk, and the difference matters when two formats overlap in the same byte range.
Once a chunk is bounded, an extractor turns it into files on disk. Extraction is recursive: the output of one pass becomes input to the next, up to a configurable depth, with a default of 10 levels. Multi-processing is on by default and uses all available CPU cores, which is why the CLI exposes --process-num for the cases where you want to cap it.
Anything that no handler claims is treated as an unknown chunk. unblob carves it out anyway and reports it, and it automatically identifies null and 0xFF padding so that padding does not dominate the report. For unknown chunks it also computes Shannon entropy and a chi-square probability, which gives you a rough signal about whether the region is encrypted, compressed or genuinely random. That entropy step is not free, so it is controlled by its own depth setting, defaulting to 1 and switchable off with 0.
The output is a directory tree plus, optionally, a JSON report carrying chunk offsets, sizes, entropy, file ownership, permissions and timestamps. The report is the part that separates unblob from a shell script built on dd: it records provenance, so a later analyst can tell which byte range produced which file.
How to install unblob and run a first extraction
The README recommends pip for most users. The Python package installs the tool itself, but not the external extractor binaries it shells out to for several formats, so the second step matters as much as the first.
pip install unblobOn Ubuntu or Debian, the README lists the external tools to install alongside it. Skipping this step is the most common way to end up with a format that unblob claims to support but cannot actually extract on your machine.
sudo apt install android-sdk-libsparse-utils e2fsprogs p7zip-full unar zlib1g-dev liblzo2-dev lzop lziprecover libhyperscan-dev zstd lz4SquashFS needs a separate package, sasquatch, which the README installs from a GitHub release. The architecture is resolved with dpkg at download time.
curl -L -o sasquatch_1.0.deb "https://github.com/onekey-sec/sasquatch/releases/download/sasquatch-v4.5.1-6/sasquatch_1.0_$(dpkg --print-architecture).deb"
sudo dpkg -i sasquatch_1.0.deb && rm sasquatch_1.0.debBefore extracting anything, check what unblob can actually see. The command prints the external dependencies and whether each one is available.
unblob --show-external-dependenciesIf you would rather not assemble that list, the Docker image bundles every extractor. The README gives this invocation, mounting an input and an output directory and passing the file inside the container.
docker run \
--rm \
--pull always \
-v /path/to/extract-dir:/data/output \
-v /path/to/files:/data/input \
ghcr.io/onekey-sec/unblob:latest /data/input/firmware.binA first real run is one command. Output goes to firmware.bin_extract/ by default, and adding --report writes the JSON metadata alongside it.
unblob --report report.json firmware.binFor anything you plan to script, the Python API exposes the same options through ExtractionConfig. The README shows randomness_depth, which is the entropy setting, and notes that max_depth, process_num, skip_magic, force_extract and keep_extracted_chunks are also accepted.
from pathlib import Path
from unblob.processing import ExtractionConfig, process_file
config = ExtractionConfig(
extract_root=Path("/tmp/output"),
randomness_depth=1,
)
result = process_file(config, Path("firmware.bin"))The external extractor dependency is the real failure mode
unblob is a Python package with a Rust component, and pyproject.toml pins a long list of Python dependencies including pyperscan, which needs libhyperscan, and lz4, which carries an explicit exclusion for version 4.4.3 because that release had no aarch64 wheels. That is a packaging problem you inherit, not one unblob can solve for you.
The larger issue is the split between formats unblob understands natively and formats it delegates. The README lists 78+ supported formats and then tells pip users to install a dozen system packages, plus sasquatch for SquashFS. If a delegated tool is missing, the format is not extracted; the bytes fall through to the unknown-chunk path. You get a carved region and an entropy figure instead of a mounted file system, and nothing in the output screams that a dependency was absent. The --show-external-dependencies command exists precisely because this is easy to get wrong, and it is worth running as a preflight check rather than after an extraction looks thin.
Recursion depth is the other boundary. The default of 10 is generous, but deeply nested firmware can exceed it, and when it does the extraction stops at that level. Raising -d costs time and disk, and the README does not document a rollback for a partially completed run, so a failed deep extraction leaves you deciding by hand what to delete.
Finally, unblob is the wrong tool when the format is proprietary and undocumented. Handlers are written against format specifications, and a vendor container with no published layout will be carved as an unknown chunk. That is the honest outcome, but it means unblob cannot do what a differential or heuristic approach might.
unblob against binwalk: bounding chunks versus scanning for signatures
The comparison people actually search for is unblob versus binwalk, and the difference is architectural rather than a matter of speed.
binwalk scans a file for magic signatures and reports where each one appears. It is excellent at telling you that a SquashFS header exists at offset 0x200000. It is weaker at telling you where that file system ends, because a signature scan has no notion of a chunk boundary unless a plugin computes one. Overlapping false positives inside compressed data are a known consequence of that approach.
unblob inverts the emphasis. Handlers compute both ends of a chunk according to the format standard, and extraction is recursive by construction, so a container inside a container is handled by the same pass rather than by re-running the tool on carved output. The unknown-chunk carving and entropy analysis are there to give a defined answer for the regions where no handler applies, rather than leaving a gap in the report.
The trade-off is dependency weight. binwalk is a single tool you can drop onto a machine; unblob asks for a Python environment, a set of system extractors, and in the source case a Rust toolchain. If you want a quick signature listing on a throwaway VM, binwalk is the lighter answer. If you want a reproducible extraction with offsets recorded, unblob's model is the one that scales to a corpus of images.
Licence, packaging and what maintenance costs you
unblob is MIT licensed, and pyproject.toml declares the same. MIT is permissive: you can use it commercially, modify it and redistribute it, provided the copyright notice and permission notice travel with it. The repository also carries a NOTICE file, which is worth reading before you vendor the code, though nothing here constitutes legal advice and the dependency licences are your own problem to audit.
That last point deserves emphasis. unblob's Python dependencies include cryptography, lief, rarfile and others, each with its own licence, and the external extractor tools installed through apt carry theirs. An MIT licence on unblob does not make the assembled toolchain MIT.
On maintenance, the repository is not archived and the last push was on 2026-06-04, which is more than three months before today but recent enough that the project is not dormant. Releases follow a calendar-style versioning scheme: 26.6.4 on 2026-06-04, 26.3.30 on 2026-03-30, 26.3.24 on 2026-03-25. The gap between the March releases and June suggests a cadence of a few releases per quarter rather than continuous churn, and the README does not describe a deprecation policy or a support window for older versions. Upgrading means re-checking the external tool list, since a new format handler can add a new system dependency that pip will not install for you.
The plugin system softens the upgrade cost for custom formats. Handlers and extractors can live outside the tree and be loaded at runtime with --plugins-path, so an internal format does not have to be carried as a patch against upstream.
Editorial conclusion
Adopt unblob if you routinely open firmware images and want chunk offsets, entropy figures and a JSON report rather than a directory of files with no provenance. Skip it if your target format has no handler and you cannot write one, or if you need extraction to run without the external tools that the pip install leaves you to install yourself. Before committing, run unblob --show-external-dependencies on the machine that will do the work and confirm that sasquatch and the file-system tools are present, because a missing extractor turns a supported format into an unrecognised chunk.
Frequently asked questions
How to install unblob?
The README recommends pip install unblob, then installing the external extractor tools with apt on Ubuntu or Debian, plus sasquatch for SquashFS. Alternatively the Docker image bundles all extractors, and Kali, Nix and a source build with uv are also documented.
how to install unblob
Same route: pip install unblob for the Python package, then the apt package list from the README, then unblob --show-external-dependencies to confirm the extractors are visible. If you would rather skip that, the ghcr.io/onekey-sec/unblob image includes them.
Does unblob need root privileges to extract firmware?
No. The README states that unblob requires no elevated privileges and runs safely as a regular user. The Docker image also creates a non-root unblob user and runs as it.
Why does unblob report a format as an unknown chunk instead of extracting it?
Most likely the external extractor tool for that format is not installed, since the pip package does not bring the system binaries with it. Run unblob --show-external-dependencies to check, or use the Docker image, which bundles them.
Can unblob extract a format that has no handler?
No. Formats are recognised by handlers written against their specifications, and a proprietary container with no published layout will be carved out as an unknown chunk and reported with entropy figures rather than extracted. The plugin system lets you add a handler for it.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/onekey-sec-unblob)