google/magika: AI File Type Detection for Gmail-Scale Pipelines
Fast and accurate AI powered file content types detection
At a glance
- What is it?
- Magika is Google's open source file content type detector, shipped as a Rust CLI, a Python package, a JavaScript package and Go bindings. It trades exact byte inspection for a small neural model that returns a label, a MIME type and a confidence score in about 5ms per file.
- Who is it for?
- Adopt Magika when you need to route files to the right scanner or parser before you open them, especially for text formats where extension and magic-byte checks disagree. Skip it if your policy requires a deterministic, spec-defined answer for every input, because a confidence threshold can downgrade a detection to a generic label.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Magika solves is misrouted files, not unknown extensions
Every ingestion pipeline that accepts untrusted files eventually needs one decision: which parser or scanner gets this file. The usual tools answer with an extension or a magic-byte signature. Both fail in predictable ways. A file named report.pdf can be a ZIP archive, a text file, or a Python script, and the first bytes of a text format rarely contain a fixed signature at all. That is where Magika positions itself. The README describes it as an AI-powered file type detection tool that relies on deep learning, and states the target: 200+ content types covering both binary and textual formats.
The audience is narrow and specific. Magika is for teams running content policy or security scanners over user uploads, where a misclassified file means the wrong scanner sees it. The README says Magika is used at scale to route Gmail, Drive and Safe Browsing files to the proper security and content policy scanners, and that it has been integrated with VirusTotal and abuse.ch. Those are routing problems, not forensic identification problems. If you need to prove what a file is to a court or a compliance auditor, this is the wrong tool.
What the model actually reads, and where the threshold sits
Magika does not read the whole file. The README states that inference time is near-constant regardless of file size because Magika only uses a limited subset of the file's content. The model itself is small: the README puts it at about a few MBs, running on a single CPU at roughly 5ms per file once loaded. Loading the model is a one-off cost, so batch use is where the per-file number matters.
The interesting design decision is the per-content-type threshold system. Rather than always returning the top class, Magika decides whether to trust the prediction for that content type. When it does not, it returns a generic label such as "Generic text document" or "Unknown binary data". That is a deliberate trade: you get fewer confident answers, but the answers you do get are meant to be safe to act on. The tolerance is adjustable through prediction modes named high-confidence, medium-confidence and best-guess. A pipeline that cannot handle a generic label is a pipeline that will break on real uploads, because the generic label is a normal output, not an error condition.
The JSON output exposes both sides of this. Alongside the resolved label, MIME type, extensions and group, the result carries a score, and the example in the README shows a value of 0.996999979019165 for a Python file. The score is the number to log and threshold on in your own code if the built-in modes do not match your risk appetite.
Installing the CLI and labelling your first directory
The CLI is written in Rust and the README lists several install paths. If you already use Python tooling, pipx is the shortest route, and it installs the CLI from the magika Python package.
pipx install magikaOn macOS or Linux, Homebrew is an alternative. There is also a Rust package, magika-cli, installable through cargo.
brew install magika
cargo install --locked magika-cliThe README also documents an installer script for POSIX shells and a PowerShell equivalent, both fetched from securityresearch.google. Once installed, the quickest real use is a recursive scan. The README's own example changes into a test data directory and passes -r with a glob.
cd tests_data/basic && magika -r * | headWhat you should see is one line per file in the form path: description (group). The README's sample output includes lines such as "c/code.c: C source (code)" and "empty/empty_file: Empty file (inode)". Note the inode group for the empty file: Magika has labels for inputs that are not really content.
For machine consumption, add --json. The README shows a single-file invocation returning an array of objects, each with a path and a result containing status, a value with both a dl and an output object, and a score.
magika ./tests_data/basic/python/code.py --jsonThe two objects, dl and output, are worth noticing. The README's example shows them identical for a confident Python detection, which suggests output reflects any post-processing applied on top of the raw model decision. If you build against this schema, read the documentation for the difference rather than assuming they always match.
Where Magika gives you a worse answer than a signature check
The threshold system is the main limitation and it is structural. A signature-based tool such as libmagic, or the file command built on it, will tell you that a file starts with the bytes for a ZIP archive. Magika may tell you it is a Microsoft Word 2007+ document, because a .docx is a ZIP container and the model has learned the surrounding structure. For routing to a document pipeline that is the better answer. For a tool that needs to know the container format, it is a worse one, and the README does not describe any way to get both answers from a single call.
There is a second boundary. The README does not document rollback, versioning of the model beyond the CLI and Python release tags, or how a model upgrade changes labels for files you already classified. If you persist Magika labels in a database, a model change is a data migration you have to plan for, and the README is silent on how to detect that a label's meaning shifted.
Finally, the Go bindings are marked WIP in the README, and the JavaScript package is described as experimental. Treat those as preview surfaces. The Python package and the Rust CLI are the paths the README presents as the main ones, and the release list shows a Python release at python-v1.0.2 and CLI releases at cli/v1.1.0.
Magika against libmagic and against a full content inspection step
The obvious alternative is libmagic, the engine behind the file command. The difference in approach is stark. Libmagic evaluates an ordered rule database against the file's leading bytes and returns the first rule that matches; it is deterministic and explainable, and you can read the rule that fired. Magika runs a neural classifier over a limited subset of content and returns a probability-like score, then applies a per-type threshold. Libmagic will happily call a file by its container format. Magika aims at the format a user or a scanner would name.
Both are cheap. Libmagic reads a header; Magika's README claims about 5ms per file on a single CPU after model load, with near-constant time regardless of file size. The practical difference is not cost, it is failure shape. Libmagic fails by matching a too-general rule. Magika fails by returning a generic label, which is at least an explicit signal that the model declined to commit.
A third option is running a full parser and catching its exception. That is more expensive, and it tells you what the file is only after you have handed it to code that may itself be the attack surface. Magika's value is that it runs before that step.
Licence, release cadence and what an upgrade costs you
Magika is Apache-2.0. That is a permissive licence with an explicit patent grant, and it is compatible with commercial use, but this is not legal advice and you should confirm obligations with your own counsel, particularly the notice and attribution requirements if you redistribute the model or the bindings.
The repository is not archived, and the last push was on 2026-09-10. The recent release tags show cli/v1.1.0 and cli-latest on 2026-04-24, and python-v1.0.2 on 2026-02-27. Those are the versions to pin against if you need reproducibility. The README does not describe a deprecation policy for model versions or a compatibility guarantee between the CLI and the Python package, so an upgrade is a change you should test against a labelled corpus of your own files rather than assume is behaviour-preserving.
The Dockerfile in the repository pins the Python package with an ARG default of MAGIKA_VERSION=1.0.3 and installs with pip install --no-cache-dir magika==${MAGIKA_VERSION}. It runs as a non-root magika user and defines a healthcheck that calls magika --version. That is a reasonable template if you want a containerised CLI, and it shows the maintainers' own pinning convention.
Editorial conclusion
Adopt Magika when you need to route files to the right scanner or parser before you open them, especially for text formats where extension and magic-byte checks disagree. Skip it if your policy requires a deterministic, spec-defined answer for every input, because a confidence threshold can downgrade a detection to a generic label. Verify first that the returned label set covers your formats and that your pipeline handles 'Generic text document' and 'Unknown binary data' as first-class outcomes rather than errors.
Frequently asked questions
What is Magika and what does it do?
Magika is an AI-powered file type detection tool from Google that uses a custom deep learning model to identify file content types. The README states it was trained and evaluated on a dataset of about 100M samples across 200+ content types and achieves roughly 99% average accuracy on the test set.
How do I install Magika?
The README lists several routes: pipx install magika for the CLI through the Python package, brew install magika on macOS and Linux, an installer script fetched from securityresearch.google for POSIX shells and PowerShell, and cargo install --locked magika-cli for the Rust package. The Python API installs with pip install magika and the JavaScript package with npm install magika.
How fast is Magika and how much memory does the model need?
The README states the model weighs about a few MBs and that, after the one-off model load, inference takes about 5ms per file even on a single CPU. It also states inference time is near-constant regardless of file size because only a limited subset of the file's content is used.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/google-magika)