CIRCL/AIL-framework: threat intelligence built around running new rules against old data
AIL framework - Analysis Information Leak framework. Project moved to https://github.com/ail-project
At a glance
- What is it?
- AIL is an AGPL licensed platform from CIRCL for collecting, crawling, processing and correlating unstructured information, and the feature that justifies its storage appetite is retro-hunting: writing a detection rule today and running it against everything collected months ago. The repository you land on is a redirect shell, which is why its release history stops in 2020 while the readme describes version 6.7.
- Who is it for?
- AIL is worth evaluating if you already collect data from the open web, hidden services or chat platforms and have discovered that your detection rules only ever fire forward, because retro-hunting is the one capability here that compounds with the size of your archive.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
You are looking at a redirect, and the version numbers prove it
The repository carries the name of the organisation that wrote the software, and almost everything inside it points somewhere else. The project description says the project has moved. The homepage field points at a different organisation. The badge for the latest release, the badge for continuous integration, the chat badge and the contributors badge all reference the new location. The logo is loaded from the new location. The default branch here receives commits as recently as mid-September 2026, so this is not a frozen copy, it is a shell that is still being touched.
The release history is where the redirection shows. The three most recent releases visible here are from early 2020: a version 3.1 with new crawling capability, a version 3.0 with full format import and export to a threat intelligence platform, and a version 2.9 which the release note describes as carrying a critical security fix with an assigned identifier. Meanwhile the readme says the software is now at version 6.7 and describes a substantial list of changes in that release: a unified search interface, date filtering, image descriptions, wider optical character recognition, a full document processing pipeline with translation, a second anonymised network, passive correlation, and better chat exploration.
All of that is consistent with one explanation: the tags were never migrated. Whoever cut version 6.7 did it in the new repository, so anyone reading the release list here is reading a history that stopped when the project changed name. For an analyst deciding whether to deploy this, that distinction matters more than it sounds. There is no tag in this repository corresponding to the version the readme describes, so there is no way to check out a specific known-good release from here, and the security fix visible in the 2.9 note is the only release note in this repository that tells you anything about a vulnerability at all.
The licence is the strongest one in the copyleft family for a network service, which is the correct choice for software that runs on a server and processes other people's data. A security file and a how-to file are both present, which is more than most projects of this size offer.
Five stages, and the framework hides its real shape in the third one
The readme organises the capability set as an intelligence lifecycle with five stages, and reading it in order is the fastest way to understand what this platform is for.
Collection is continuous ingestion from chat platforms, websites, hidden services on both of the two anonymising networks it supports, files, and external feeds. There is a separate crawler manager described as handling continuous and authenticated collection, and the authentication is the interesting part: it reuses browser sessions, cookies and local storage, which means the crawler can log in as a user rather than only fetching public pages.
Processing is where most of the dependency list ends up. It extracts addresses, hostnames, email addresses and credentials, detects phone numbers, API keys, international bank numbers, certificates and private keys, and recognises cryptocurrency artefacts including addresses and private keys. It decodes encoded file content, runs optical character recognition on screenshots, parses QR codes and barcodes and reprocesses what they contain, generates descriptions of images and domains, and extracts and translates document text and metadata. Tagging draws on two shared taxonomies maintained by the same ecosystem as the export target.
Detection is the stage the platform is actually organised around, and it is where the two ideas that make it interesting live. Analysis is search, pivoting and correlation graphs. Dissemination is export to threat intelligence sharing platforms, and the format support is bidirectional.
What the five-stage framing hides is that stages one, two and four are table stakes and stage three is the product. Anyone can collect a feed and index it. What makes this platform worth its storage cost is that stage three is a rule engine you can point backwards in time, and that is the next section.
Retro-hunting is the argument for keeping everything
A tracker in this platform is a user-defined rule or pattern that detects, tags and can notify an analyst about something in the incoming stream. The supported types are a single word, a set of words, a regular expression, a YARA rule, and typo-squatting detection.
On its own that is a familiar alerting feature, and it has the familiar weakness: it only sees what arrives after you write it. The rule you write in March says nothing about the two years of material already sitting in the archive, which is where the interesting material usually is, because a leak is a leak whether or not anyone was looking.
The feature that answers this is called a retro hunt, and the readme describes it as enabling analysts to run newly created rules against historical data to uncover previously missed content. That single capability changes the economics of the whole platform. It converts the archive from a passive record into a repeatedly re-derivable asset, and it changes the value of a detection rule from a one-off alert to a permanent addition to what you can ask about your own data.
There is a YARA editor built in, which matters for the same reason. Writing a binary signature by hand in a text editor is unpleasant enough that people do not do it, and a rule you will not write is a rule you will not have when the retro hunt runs.
The correlation engine is the other half of the retrospective value. The readme lists the relationship types it can graph: decoded files against their hashes, pretty good privacy metadata, domains against titles, a domain hash, favicons and cookie names, usernames against accounts, vulnerability identifiers, SSH keys, cryptocurrencies and document metadata. A correlation graph over two years of archives is a different artefact from a correlation graph over a week, and it is the reason the platform stores what it stores rather than sampling.
For an evaluator, the question to ask is therefore not how fast the ingest is but how much history you already have and can bring with you. If the answer is none, the retro hunt has nothing to chew on and the platform is an expensive alerting tool.
The dependency list is the architecture, and it is unusually honest
The requirement file is grouped by purpose with comments, which is rarer than it should be and makes this the most informative file in the repository. Reading it tells you what the platform actually does, and several entries explain a capability that the readme states without justification.
The integration layer is four lines and it explains the whole dissemination story: a first-party client, a search client, the threat intelligence platform library with a version floor, a client for a second case management platform, and two taxonomies installed from source. Both taxonomies are pulled as raw repository URLs, which is a detail that returns in a moment.
The core is a cache server with a version range, file type identification, YARA bindings, a regular expression engine with a Rust implementation, and a fast serialiser. Then there is a block of three hash implementations, which tells you something about the correlation graph: a CRC, a non-cryptographic hash used widely in feed deduplication, and a fast non-cryptographic hash. Choosing three means matching hashes that arrived from different sources.
A messaging library and its Python binding appear next, and a natural language toolkit with tokenisation and text processing. Language handling is a first-class concern, and the evidence is a Google language detection binding, a translation client, and a language detection package. Emoji handling is a fork maintained by the same project, installed from source, which is the kind of entry that tells you upstream would not take a change.
Graphics are a numerical library, a plotting library and a graph library, which is the correlation visualisation. System information comes from a terminal user interface library and a process library, and the terminal interface entry implies an operator console as well as a web one.
Then the specialised blocks, each matching a readme bullet. A domain reputation ranking package and a domain classifier for the domain analysis features. A typo-squatting package and a Levenshtein distance library for the lookalike detection. A phone number library. Optical character recognition, plus two separate QR and barcode readers, one of which needs a native system library. Two document libraries, one of them the machine-learning-oriented variant of the other, for the pipeline that now includes translation.
The crawler block is a scraping framework plus a rendering integration, which is how it gets JavaScript pages and authenticated sessions. And the web tier is a microframework, a session login extension, a socket extension, a password hashing library, a one-time password library, and a QR code generation library. Password hashing and one-time passwords in the same block is the answer to whether the platform's own users are handled properly, and they are.
The test stack is a runner plus a coverage tool, and the runner is the second-generation successor to the original nose framework. Old, but consistent with a codebase of this vintage.
Three dependencies have no version, which is the build problem
Two entries in the requirement file are raw repository URLs rather than versioned releases: the two taxonomies the tagging system draws on, and the project's own emoji fork. A third block installs a validation binding that is commented out, which reads as a feature that was trialled and set aside rather than one that was removed.
Installing from a branch means your build resolves to whatever the tip of that repository is on the day you run it. For a platform this size, where the taxonomies are load-bearing for tagging and the emoji library is used in processing, that is the difference between a deployment you can rebuild in two years and one you cannot. It is the most consequential line in the file and there is no comment explaining it.
By contrast, one dependency is pinned at both ends, with a floor and a ceiling around the cache server client. A bounded range like that is a deliberate statement that versions above the ceiling are known to break something, which is exactly the right way to handle a component whose client protocol moves. It is also the only line in the file that shows this kind of care, and the contrast between that one entry and the two unversioned repository URLs is the honest summary of the dependency policy.
The tail of the file contains a section of retired packages, all commented out: a text table library, and three download URLs for address lookup tooling, two of which point at archive services for code that no longer exists in its original home. That section is harmless and also informative. It is what a decade-old requirement file looks like when somebody has been keeping it tidy rather than deleting it, and it is a warning to anyone who plans to re-enable anything from it.
The practical advice follows directly. Build inside a virtual environment, which the repository provides a script for, and capture the resolved set of packages on the day you install, because the file will not do it for you. The install scripts and the how-to file are where that guidance lives; the readme is not.
A search server, a cache, a collector manager, and one feature to read carefully
Three architectural facts are worth pulling out of the dependency list, because they determine what you have to operate.
The indexer is a search engine written in Rust with its own server process, not a document store you embed. That is a real choice and it has consequences: you get fast typo-tolerant search without writing a query language, and you take on a second service to run, upgrade and back up. The readme's emphasis on a unified search interface with best-match and most-recent ordering is the payoff, and it is a genuine one, because most homegrown search across chat archives is where projects like this fall down.
The cache server is already implied by the bounded client version, so that is a second service. The collector manager is a third, and it is the one that carries the authenticated crawling, which is the platform's most operationally demanding feature: it maintains browser sessions, cookies and local storage, which means it holds credentials for whatever sites you have logged into. Whatever the rest of the platform does, that component needs a threat model of its own, and the readme does not provide one.
The feature to read carefully is passive correlation, which the readme lists among the recent additions and describes as being for infrastructure analysis and deanonymisation workflows. Passive means it observes connections the platform's own collection already touches, which is a meaningfully different thing from active probing, and that distinction is why it is in a defensive platform rather than a red-team toolkit. Even so, of everything on the feature list this is the one with the highest dual-use weight and the least operational justification offered. If you are evaluating the platform for an organisation, that is the line to ask a question about, because the answer determines whether the deployment needs a review process around it.
Set against that, the rest of the list is squarely defensive. Cookie names, favicons, domain hashes, SSH key fingerprints and document metadata are all things a defender correlates to find infrastructure they did not know they had.
Editorial conclusion
AIL is worth evaluating if you already collect data from the open web, hidden services or chat platforms and have discovered that your detection rules only ever fire forward, because retro-hunting is the one capability here that compounds with the size of your archive. It is a poor fit as a quick install, since three of its dependencies are pulled straight from source control with no version constraint, one search dependency needs a separate server, and the whole thing expects a virtual environment plus a managed cache, so a reproducible build is not something you get for free. Before deploying it, read the installation scripts and the how-to rather than the readme, and read the passive correlation feature carefully, since that is the part of the platform with the most dual-use weight and the least operational justification documented.
Frequently asked questions
What is the AIL framework used for?
It is an open-source platform for collecting, crawling, processing and analysing unstructured information from the clear web, hidden services on two anonymising networks, chat platforms, files and external feeds, to support threat intelligence and leak analysis. It was originally developed at CIRCL and its current home is a different organisation.
What is a retro hunt in AIL framework?
It is the ability to run a newly written detection rule against historical data already collected, to find content that was missed when the rule did not yet exist. The readme presents it as the way analysts get value from an archive of previously ingested material.
Which detection rule types does AIL framework support?
Trackers can be a single word, a set of words, a regular expression, a YARA rule, or a typo-squatting detector. There is a built-in YARA editor, and trackers can tag objects and trigger webhook or email notifications.
Why does the AIL framework release history stop at version 3.1 when the readme describes version 6.7?
Because the project moved to a different organisation and the release tags were not migrated. The most recent tags visible in this repository are from 2020, while the readme describes a much later version, so there is no tag here matching the version the documentation describes.
What services does running AIL framework require?
At least a cache server, a separate search index server, and the collector manager that handles crawling including authenticated sessions. The web interface is built on a Python microframework with session login, password hashing and one-time password support, and the platform integrates with threat intelligence and case management platforms for import and export.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/circl-ail-framework)