AIL's readme says version 6.7 while the newest release is 7.1
AIL framework - Analysis Information Leak framework
At a glance
- What is it?
- A leak-analysis and threat-intelligence platform developed at a research centre, collecting from the clear web, Tor and I2P, with YARA retro-hunts over historical data. Its dependency file caps Redis, installs four packages from unpinned git URLs, and keeps a section of dead archive links as comments.
- Who is it for?
- AIL fits an analyst team that already has collection problems and wants one pipeline for extraction, detection and sharing rather than five tools. The parts that make it more than a search box are the retro-hunt over historical data, the correlation graph across decoded files, hashes, PGP metadata, domains, cookies, keys and CVEs, and the export path into MISP with its galaxy and taxonomy vocabulary.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The readme announces 6.7 and the repository is on 7.1
The features section has a heading about what is new, and the first line under it states the current version. That line says 6.7.
The release list says otherwise. The newest tag is 7.1, published in late September 2026, and the one before it is 7.0, published in early July. A third recent tag is 6.9 from June. So the readme is two minor versions behind the tags, and it is behind in the one section a reader is most likely to trust: the changelog summary.
The release titles themselves are informative about where the work has gone. The 7.0 title is about forums, crawling and collection, and 7.1 is about interactive crawling, forum improvements and image similarity. Two consecutive releases on forums and crawling suggests a sustained effort in one area rather than a scatter of features.
What is missing from the readme is any account of 7.0 and 7.1 beyond the headline. Interactive crawling in particular is the kind of change that alters how the tool behaves against a site, and the visible documentation does not say what it now does differently.
For a project on a weekly release cadence, a version line in the readme is going to rot. The release list is the source of truth.
Redis has a ceiling and six requirements have no version at all
The dependency file is organised into labelled blocks, which makes it easy to audit, and the pinning is inconsistent in ways worth knowing.
One requirement carries an upper bound, which is rare and consequential:
redis>4.4.4,<=7.4.1Capping the Redis client at a major version below eight means this file will eventually refuse to resolve, and it will do so with a resolver error rather than a clear message. Everything else with a constraint is a floor. If the cap exists because of a real incompatibility, that incompatibility belongs in a comment; if it is caution, it will become a blocker.
Then there are the entries with no version at all. A serialisation library, a message queue binding and its Python wrapper, two PDF libraries, the search index client, two of the three QR code readers, a country database, a domain age library, the HTTP client and an encoding detector are all written as bare names.
Two of the bare names are the project's own packages, published without an organisation prefix, alongside two more of its packages installed from git. So a third of the unpinned surface is first-party, which means the same author controls the version and nothing pins it.
Four requirements install from git with no reference
Four lines in the dependency file are direct git addresses rather than published packages:
git+https://github.com/MISP/PyTaxonomies
git+https://github.com/MISP/PyMISPGalaxies
git+https://github.com/ail-project/demojiNo branch, no tag, no commit. Each of those installs whatever the default branch pointed at when the install ran, so two installations a month apart can produce different taxonomies, different galaxy clusters and a different emoji extraction library.
Two of the three are the vocabularies that AIL's tagging depends on. MISP taxonomies and MISP galaxies are not decorative here: the tagging system is described as using them, and the export path writes in MISP formats. So the classification vocabulary is unpinned, which means two analysts running the same version of AIL can tag the same document differently without either of them doing anything wrong.
That is a defensible choice for tracking upstream, and it is the wrong default for a forensic tool where reproducibility is the point. The fix is to pin those three to commit hashes, which the tooling supports and which costs nothing at install time.
A whole section of the dependency file is comments
Near the bottom of the file there is a heading for packages that have been retired, and everything under it is commented out. It records two abandoned ASN lookup dependencies, one of them a tarball hosted on a Google Code archive that has been shut down for years, and two others as direct GitHub zip addresses.
Leaving the record is good practice. A reader who wonders why domain reputation lookups are not in the file can see that they were tried, and a future maintainer restoring them would find the old addresses rather than starting from scratch.
Higher up there is a similar case. A heading for bindings to a SQL injection detection library, with the single requirement beneath it commented out. That one tells you a detection capability was considered and is not currently wired up, which is more informative than a missing heading would be.
Taken together, the two are a small maintenance diary inside the manifest. They also add to the length of a file that now has well over sixty entries, which is a lot of surface for a single install step.
Three language detectors and a rendering service nobody mentions
The dependency file carries three separate language identification libraries in one block: a compact Google detector binding, a Levenshtein distance library, and a small language detector. Three overlapping tools for the same job is either a staged migration or three people solving the same problem, and nothing in the file says which.
The translation client sits in the same block, which matches the documented support for translating PDFs and crawled text.
The crawler block is the one with a hidden system dependency. Scraping is handled by a well known framework, and alongside it is a splash renderer binding, which is a client for a separate JavaScript rendering service that has to be running somewhere. Nothing in the visible installation instructions mentions installing or running that service.
That matters for two documented features. Interactive and authenticated crawling, which is how AIL handles pages that need JavaScript and pages behind a login, both depend on rendering that a plain HTTP client cannot do. A deployment that installs the dependencies and starts the web interface will find those features quietly doing less than the feature list claims.
The OCR and QR blocks add two more QR readers and an OCR library, which is consistent with the screenshots and QR extraction features and with the perceptual hashing block beside them.
Authenticated crawling replays stored sessions, including local storage
The collection section lists authenticated crawling with browser sessions, cookie reuse and local storage reuse, and the screenshot captions name the technique plainly: login-protected crawling with pre-recorded session cookies.
That is the most security-relevant capability in the tool, and it is worth being precise about what it means. AIL is not logging in as itself. It is replaying a session captured from a real browser, which means whatever that session can reach, AIL can reach, with the same permissions, until the session expires or is revoked.
The consequences are practical. The session material lives in the tool's runtime data, so that directory needs the same handling as a credential store. A crawler running with an administrator session against a forum will collect what that administrator can see, not what an anonymous reader could. And because sessions expire, a long-running crawl degrades quietly rather than failing loudly.
Set against that, the retrieval scope is entirely defensive: extracting URLs, hostnames, email addresses and credentials from material an analyst already has access to, and detecting phone numbers, API keys, IBANs, certificates, private keys and cryptocurrency artefacts so they can be spotted in a leak. That is triage tooling, and the correlation graph exists to help an analyst pivot on what has already been collected.
The screenshots section is nine headings and nothing under them
There is a screenshots section in the readme, and it is a list of nine subheadings: websites, forums and hidden services, login-protected crawling with pre-recorded session cookies, extracted and decoded files, the correlation engine, investigation, the tagging system, MISP export, automatic events and alerts, and the user interface submission path.
In this copy, nothing is rendered beneath any of the nine. There is no image and no description, only the headings.
The order of the list is a reasonable tour of the product, and reading it as a table of contents tells you what a new user is meant to see first. But it is the only place the tool demonstrates itself, and the two most interesting claims in the project, session replay and correlation, have no visual evidence attached.
The documentation elsewhere points at a project site, and the rest of the readme leans on prose and feature lists, so the missing images are a documentation gap rather than a functional one. Still, for a project whose argument is that analysts should be able to follow an object's life across a collection, a correlation graph screenshot would carry more than the paragraph describing it does.
The install command ends mid-address
The installation section opens with a shell block whose first line is a clone command, and in this copy that line stops partway through the repository address, after the organisation name and the first characters of the project name.
So the one command a new user copies first is incomplete here. Everything after it in the block is gone too, which means the rest of the installation sequence, whether it is a virtual environment script, a dependency script or something else, is not visible in this copy either.
What the repository tree does tell you is that installation is script-driven rather than a single command. There is a script that creates a virtual environment, a second that installs dependencies, a third that resets the installation, and a directory holding installers for other environments. There is also a how-to document, an agent instructions file, a security policy, a changelog configuration, and a set of runtime directories for files, logs and variable data that are part of the working tree layout rather than something created on first run.
There is no container recipe at the top level, so a container deployment is not the documented default even though container deployment is the first thing most people will look for.
Editorial conclusion
AIL fits an analyst team that already has collection problems and wants one pipeline for extraction, detection and sharing rather than five tools. The parts that make it more than a search box are the retro-hunt over historical data, the correlation graph across decoded files, hashes, PGP metadata, domains, cookies, keys and CVEs, and the export path into MISP with its galaxy and taxonomy vocabulary. Two things to weigh before deploying it. It is AGPL, and it reuses stored browser sessions including local storage to crawl authenticated sites, so treat its runtime data directory as holding credentials-adjacent material and its licence as a decision your legal team should see. And the pinned dependency set is loose in ways that will bite you: four requirements install from git with no ref, so two installs a month apart can differ.
Frequently asked questions
What is the AIL framework?
An open-source platform for collecting, crawling, processing and analysing unstructured information from the clear web, Tor hidden services, I2P, chats, files and external feeds, licensed AGPL-3.0 and originally developed at a European research centre. It covers an intelligence lifecycle from collection through processing, detection and analysis to dissemination into MISP.
What kinds of trackers does the AIL framework support?
Five: word tracking, set-of-words tracking, regex tracking, YARA rules and typo-squatting detection. Trackers automatically detect, tag and notify analysts, and the framework adds real-time tagging, object occurrence tracking, webhook or email notification, and a built-in YARA editor.
What are Retro Hunts in the AIL framework?
A feature that runs newly written YARA rules against data already collected, so content that earlier rules missed can be surfaced retroactively. That is what turns the YARA support from a live filter into a way to test a new indicator against a historical collection.
Which platforms does the AIL framework integrate with?
MISP is the integration named throughout: alerting and sharing to it, export of objects and investigations in MISP formats, and automatic exports on selected detections and tags. Its tagging vocabulary follows MISP Galaxy clusters and MISP taxonomies, and the dependency file also includes a client for a collaborative incident-response platform.
How do I install the AIL framework?
The installation section opens with a git clone command, which in this copy ends partway through the repository address. The repository tree shows the rest of the approach: a script that creates a virtual environment, a script that installs dependencies, a reset script, a directory of installers for other environments, and a how-to document, with no container recipe at the top level.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ail-project-ail-framework)