Library / SDK
horsicq/Detect-It-Easy avatar
horsicq/Detect-It-Easy

Detect It Easy: a signature and heuristic file identifier for reverse engineers

Program for determining types of files for Windows, Linux and MacOS.

11,578 stars961 forksJavaScriptMIT

At a glance

What is it?
Detect It Easy (DiE) combines native format parsers with an extensible DiE-JS analysis layer to identify packers, protectors and toolchains across Windows, Linux and MacOS. This article covers what it detects, how the heuristic engine works, how to install it, and where it is the wrong tool.
Who is it for?
Detect It Easy is aimed at malware analysts, reverse engineers and forensic examiners who need a fast, scriptable first pass over an unknown binary, and it is also useful to anyone who just wants to know what a file is. It is not a sandbox and not a full disassembler, so it will not execute a sample or reconstruct a program for you, and it will not replace a debugger when you need to observe runtime behaviour.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 8 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Detect It Easy answers that file(1) does not

Run file on a packed Windows executable and you get one line: PE32 executable. Detect It Easy is built for the next question, which is what produced this binary and what was done to it. The README describes it as a cross-platform tool for file identification and static inspection used by malware analysts, cybersecurity experts and reverse engineers, supporting compact signatures and multi-stage static analysis across Windows, Linux and MacOS. The audience is narrow and specific. Someone triaging a suspicious sample wants to know whether it is packed with a commercial protector, whether it was built with MSVC or MinGW, whether it carries .NET metadata, and whether its entry point looks synthetic. DiE is designed to answer those questions without launching the file. That last point is the design constraint that shapes everything else in the project: the analysis is static, so it can run on a sample that would be dangerous or impossible to execute.

Native parsers plus a JavaScript rule layer

The architecture is a split. A native core handles the parts that need to touch bytes efficiently: format parsing, bounded reads, address translation, searching and disassembly primitives. On top of that sits DiE-JS, the implementation language of the analysis modules, which the README is careful to distinguish from a few native detector calls. Rules live in the db/ tree as .sg files, and the repository also carries db_extra/, db_custom/, dbs_min/, dbs_special/, peid_rules/ and yara_rules/ directories, so the detection database is a text artifact you can read, diff and extend rather than a compiled blob. The README frames this as keeping detection policy reviewable in the database. That is a real property. You can open a rule, see which structural facts it tests, and disagree with it. The PE heuristic engine is a good example of how far the JavaScript layer is pushed. It is credited in the README to DosX and lives at db/PE/__GenericHeuristicAnalysis_By_DosX.7.sg. When enabled it runs as a separate higher-level engine that can corroborate, qualify or reject earlier database results, and every heuristic conclusion is marked as such. The file is never launched; the engine works on headers, data directories, sections, imports, exports, resources, .NET metadata, bytecode, overlays, debug records and reachable startup code rooted at the entry point. It also covers DLL initialization code, which the README calls a common blind spot when suspicious behaviour begins inside a library. For native code the engine combines cached linear disassembly with bounded traversal and purpose-built state machines, tracking register, flag, stack, address-provenance and instruction-boundary facts. The README is explicit that it is not a sandbox and not a full CPU emulator, and that it stays effective across register substitution, neutral padding, equivalent arithmetic forms and bounded reordering. For managed code it uses an internal MSIL opcode model to build operand-aware bytecode patterns, again without being a CLR emulator.

Installing Detect It Easy and running a first scan

The README points to the releases page at github.com/horsicq/DIE-engine/releases for downloads, and it carries a warning worth repeating: detectiteasy.com is not affiliated with the project or its official website, and should not be used to download DiE. The repository ships a Dockerfile that installs the Ubuntu 24.04 amd64 Debian package, which is the most reproducible route on Linux. It takes the version as a build argument, so you can pin it:

dockerfile
FROM ubuntu:24.04
ARG DIE_VERSION=4.0.0
RUN wget https://github.com/horsicq/DIE-engine/releases/download/Beta/die_${DIE_VERSION}_Ubuntu_24.04_amd64.deb

The same file then installs the package, removes the bundled database with rm -rf /usr/lib/die/db, and copies the repository's own db directory into /usr/lib/die/db. That copy step is the part to notice. The container does not use whatever database shipped inside the .deb; it uses the one in the source tree, which is how you keep the rules in step with a specific checkout. The entry point is /usr/bin/diec, the command line build. Building the image and running it against a file gives you a text report rather than a window:

bash
docker build -t die .
docker run --rm -v "$PWD:/samples" die /samples/suspect.exe

On a desktop install you get the graphical front end instead, which is what the screenshots in docs/ show. Either way the first useful move is the same: open a file you already understand, such as a binary you compiled yourself, and read which signatures fire. That tells you more about the database than any documentation will, because you can see which of your assumptions the rules agree with.

What the heuristic engine cannot tell you

The README's own framing is the limitation. The PE heuristic engine is described as tracking the facts each rule requires without pretending to be a sandbox or a full CPU emulator. That means a packer that decrypts its payload only after several API calls, or that resolves imports at runtime through a hash, produces no evidence for the static passes to correlate. Self-modifying and position-independent stubs are listed among the things the engine can expose, but exposure is not unpacking. If you need the actual unpacked code, DiE is a pointer to the right tool, not the tool. The second limitation is the nature of heuristic output. The README says heuristic conclusions are visibly marked and that the engine can corroborate, qualify or reject earlier database results. That is a design that admits uncertainty, and it means a heuristic hit is a hypothesis with stated evidence, not a verdict. A third constraint is that the analysis is bounded by design: the README mentions bounded scans and bounded control flow as a way to control false positives. Bounded also means incomplete. Code reachable only through an indirect jump that the traversal cannot resolve will not be visited. Finally, DiE is an identification tool, not a triage platform. It has no sandbox, no network capture, no memory forensics. If your question is what a sample does when it runs, this is the wrong tool and you should reach for a debugger or an instrumented environment.

Detect It Easy compared with a YARA-only workflow

The obvious alternative for signature work is YARA, and the difference is architectural rather than a matter of rule quality. A YARA rule describes byte patterns and conditions over the raw file. It has no model of the PE structure, no address translation, no disassembly, and no notion of an entry point whose reachable code can be walked. To write a YARA rule that reasons about a section's virtual size relative to its raw size, you encode the header offsets yourself, and the rule breaks when a different toolchain lays the file out differently. Detect It Easy puts the parser underneath the rule. The README describes the native core as supplying format parsing, bounded reads, address translation, searching and disassembly primitives, and the JavaScript modules as building algorithms on those primitives. That is why the PE heuristic engine can relate version-resource identity to Authenticode state, Rich build metadata, runtime model and detected protection in one pass. A YARA rule can express that correlation too, but you are writing the parser first. The trade-off runs the other way as well. YARA is a single, widely deployed engine with rules you can lift between products, and the repository does carry a yara_rules/ directory, so the two are not mutually exclusive. If your pipeline already standardises on YARA, adding DiE means adding a second engine and a second database to keep current. If your analysts are reading files by hand, DiE's structured report is the more useful artifact.

Database releases, licences and keeping rules current

The project separates code from data. Recent releases listed on the repository include db and db_extra, both dated 2026-04-16, and an earlier current-database release from 2025-03-25. The last push to the repository was on 2026-09-21. The practical consequence is that a database update does not require rebuilding the application, and the Dockerfile shows the mechanism: delete /usr/lib/die/db and copy in a fresh tree. On a desktop install the same principle applies to whatever directory the package installs the database into. The upgrade cost is therefore mostly a matter of deciding how often you pull a new db and re-run your own samples against it, because a rule change can move a file from one verdict to another without any code change. The licence is MIT, which is permissive and places few conditions on redistribution or modification. Note the boundary: the licence covers the project as published, and the README credits individual contributors for specific modules, such as the PE heuristic engine maintained by DosX. If you plan to redistribute a modified database or bundle DiE inside a product, check the terms of the specific artifacts you ship rather than assuming a single licence covers everything in the tree. Nothing here is legal advice.

Editorial conclusion

Detect It Easy is aimed at malware analysts, reverse engineers and forensic examiners who need a fast, scriptable first pass over an unknown binary, and it is also useful to anyone who just wants to know what a file is. It is not a sandbox and not a full disassembler, so it will not execute a sample or reconstruct a program for you, and it will not replace a debugger when you need to observe runtime behaviour. Before you adopt it, verify three things: that the release you downloaded comes from the links in the horsicq/Detect-It-Easy repository rather than detectiteasy.com, which the README states is not affiliated with the project; that the database shipped with your package matches the database release you intend to use, since the Dockerfile deletes /usr/lib/die/db and copies a fresh one from the repository; and that the heuristic engine's output is treated as evidence to check, not as a verdict, because the README describes it as a set of passes that corroborate, qualify or reject earlier database results.

Frequently asked questions

What does Detect It Easy do?

It identifies file types and performs static inspection of binaries on Windows, Linux and MacOS, combining native format parsers with an extensible DiE-JS analysis layer. The README describes its users as malware analysts, cybersecurity experts and reverse engineers. It never launches the file it analyses.

How do I install Detect It Easy on Linux?

The README links to the releases page at github.com/horsicq/DIE-engine/releases. The repository also ships a Dockerfile that installs the Ubuntu 24.04 amd64 Debian package and copies the repository's db directory into /usr/lib/die/db, with /usr/bin/diec as the entry point.

How do I use Detect It Easy on a file?

With the graphical build you open the file and read the report. With the containerised build the Dockerfile sets /usr/bin/diec as the entry point, so you pass a path to the binary and get a text report. The README notes the file is never launched.

Is Detect It Easy safe to download?

The README warns that detectiteasy.com is not affiliated with the Detect It Easy project or its official website, and says not to trust it as an official source or use it to download DiE. It directs users to the release links in the repository instead.

Is Detect It Easy free?

The project is published under the MIT licence, which is permissive. The README also carries a PayPal donation badge, but the repository does not describe a paid edition.

What is a good alternative to Detect It Easy?

YARA is the closest alternative for signature work, but a YARA rule operates on raw byte patterns and conditions without a structural model of the file. Detect It Easy supplies format parsing, bounded reads, address translation, searching and disassembly primitives to the DiE-JS modules built on top of them. The repository also carries a yara_rules/ directory.

Official sources

  1. horsicq/Detect-It-Easy on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/horsicq-detect-it-easy.svg)](https://hysenlabs.com/projects/horsicq-detect-it-easy)