Library / SDK
H4ckForJob/dirmap avatar
H4ckForJob/dirmap

H4ckForJob/dirmap: A Config-File Driven Web Directory Scanner

An advanced web directory & file scanning tool that will be more powerful than DirBuster, Dirsearch, cansina, and Yu Jian.一个高级web目录、文件扫描工具,功能将会强于DirBuster、Dirsearch、cansina、御剑。

3,375 stars554 forksPythonGPL-3.0

At a glance

What is it?
dirmap is a Python web directory and file scanner that puts almost all of its behaviour in dirmap.conf rather than command-line flags. The README describes four scan modes, recursive scanning and per-domain output. Here is what the repository actually documents, and where it stops.
Who is it for?
Adopt dirmap if you want a Python scanner whose scan modes, status-code filters and request behaviour live in one dirmap.conf file you can diff and version alongside a target list, and if you are comfortable editing that file because the README states detailed configuration is not available through command-line arguments.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 15 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What dirmap is trying to replace, and for whom

dirmap is a web directory and file scanner written in Python. The README frames it against a specific set of predecessors by name: DirBuster, Dirsearch, cansina and the Chinese tool 御剑. The stated goal is to be more capable than those, and the README lists the capabilities it thinks such a tool needs: a concurrent engine, dictionary support, pure brute force, crawling to build dictionaries dynamically, fuzz scanning, custom requests and custom handling of responses. Those seven items are the design brief, and each one maps to a configuration block in dirmap.conf.

The intended user is someone doing authorised web enumeration, typically during a penetration test or an internal assessment. The input formats in the README make that clear: a single host, a CIDR block such as 192.168.1.0/24, and an address range such as 192.168.1.1-192.168.1.100. That is a tool built for sweeping an internal network segment, not for probing one public URL. If your work is a single application behind a WAF, the breadth of input handling buys you nothing.

Four scan modes behind one config file

The mechanism is a configuration-driven engine. dirmap.py parses arguments, initialises a target list, loads dirmap.conf, and then dispatches to whichever scan mode is enabled. The README is explicit that the four modes are mutually exclusive: dict mode, blast mode, crawl mode and fuzz mode, one at a time. Each is toggled by an integer in the [ScanModeHandler] block, where 0 means off and, for the dictionary and fuzz modes, 1 means a single dictionary and 2 means multiple.

Dict mode reads a wordlist from conf.dict_mode_load_single_dict, or a directory of wordlists from conf.dict_mode_load_mult_dict. Blast mode generates payloads instead of reading them, bounded by conf.blast_mode_min and conf.blast_mode_max, built from conf.blast_mode_custom_charset. Crawl mode parses responses with an XPath expression, conf.crawl_mode_parse_html, whose default is //*/@href | //*/@src | //form/@action, and can also generate dynamic payloads from a suffix wordlist. Fuzz mode inserts a label into a URL template; the default label is {dir}, so http://target.com/{dir}.php becomes one request per dictionary line.

The concurrency model is one coroutine limit per target host, set by conf.request_limit, default 30. The README describes the overall shape as n targets times n payloads, so the limit applies per host rather than globally. That distinction matters on a subnet scan: a /24 with 30 threads per host is a very different load profile from 30 threads total.

Installing dirmap and running a first scan

The README gives a single setup line that clones the repository, changes into it, and installs the Python dependencies from requirement.txt. Note the filename: it is requirement.txt, singular, not requirements.txt.

bash
git clone https://github.com/H4ckForJob/dirmap.git && cd dirmap && python3 -m pip install -r requirement.txt

After that, the quickest usable command targets one host and loads the config file. The -i flag takes the target and -lcf tells dirmap to load dirmap.conf, which the README marks as required for detailed configuration.

bash
python3 dirmap.py -i https://target.com -lcf

For a subnet, the same flag accepts CIDR notation, and for a range you write the two addresses with a hyphen. The README shows both forms.

bash
python3 dirmap.py -i 192.168.1.0/24 -lcf
python3 dirmap.py -i 192.168.1.1-192.168.1.100 -lcf

If your targets live in a file, -iF reads them, and the README states the file accepts the same formats as -i.

bash
python3 dirmap.py -iF targets.txt -lcf

Results land in an output folder at the project root, one txt file per target named after the target domain. The README says results are deduplicated automatically. Before the first real scan, open dirmap.conf and check that conf.dict_mode is set to 1 and conf.dict_mode_load_single_dict names a wordlist you actually have, since that is the mode the README ships enabled.

Recursive scanning and the false-404 check

Two behaviours are worth understanding before you trust a result set. The first is recursion. The [RecursiveScan] block has conf.recursive_scan set to 0 by default, so recursion is off unless you turn it on. When enabled, it descends only into paths that returned one of the codes in conf.recursive_status_code, default [301,403], stops when a URL exceeds conf.recursive_scan_max_url_length, default 60, and skips extensions listed in conf.recursive_blacklist_exts, which covers html, images, js, css, pdf and media. You can also exclude directories outright with conf.exclude_subdirs.

The second is false-404 detection. conf.auto_check_404_page defaults to True, and the README pairs it with conf.custom_response_page, a regular expression for matching page content. This is the part of the design that most directly addresses a real problem: on an application that returns HTTP 200 with a soft 404 body, a naive scanner reports every guessed path as a hit. dirmap's answer is a regex you write yourself. That puts the accuracy of your results on your ability to craft that pattern, and the README offers only a placeholder example, ([0-9]){3}([a-z]){3}test, with no guidance on building one for a real site.

The response side is similarly filter-driven. conf.response_status_code defaults to [200], so a scan that should surface 403s or 301s quietly drops them unless you edit that list. conf.response_header_content_type and conf.response_size default to 1 and record the content type and page size; conf.skip_size, default "None", lets you discard responses of a given size.

Where the configuration model gets in the way

The README states plainly that detailed configuration is done by loading a config file and that command-line arguments are not supported for it. That is a deliberate choice and it has costs. You cannot vary a scan parameter per run without editing dirmap.conf, which means either sed-ing the file in a wrapper script or keeping several copies. On a multi-target engagement where one host needs recursion and another does not, the file becomes the coordination point.

The README also flags a set of features as not implemented, and the honesty is useful. conf.crawl_mode_parse_robots is marked 暂未实现 (not yet implemented). conf.request_max_retries is not implemented. conf.request_persistent_connect, the session reuse option, is not implemented. conf.request_header_401_auth is not implemented, with the README noting that custom request headers can cover the case. The [PayloadHandler] section is empty and marked not implemented. conf.update, which would fetch updates from GitHub, is not implemented. The TODO list in the README confirms the same gaps from the other direction, with unchecked boxes for retry limits, persistent connections and the a-z and 0-9 blast charsets.

A scanner without retry logic behaves differently on a flaky network than on a clean one: a dropped connection is a miss, not a retry. That is a real limitation to weigh if you plan to scan over a VPN or a congested link. The README does not document rollback or resume behaviour for an interrupted run either, so a scan that dies halfway leaves you to work out what was covered.

How dirmap differs from ffuf and feroxbuster

The closest alternatives in current use are ffuf and feroxbuster, and the difference is in where configuration lives. ffuf is driven entirely by flags: you pass -w for the wordlist, -u for the URL with the FUZZ keyword, and -mc to choose which status codes count as matches. Nothing persists between runs unless you write it down. dirmap inverts that. Its scan mode, its status-code filter, its recursion rules and its request headers all live in dirmap.conf, and the command line carries only the target and the flag to load that file.

The practical consequence is reproducibility. A dirmap.conf checked into a repository next to a targets.txt is a record of exactly how a scan was configured, which is useful when you need to re-run the same enumeration months later or hand the setup to a colleague. With ffuf, that record is a shell history entry or a wrapper script. The trade-off is flexibility: changing one parameter in dirmap means editing a file, and running two different configurations at once means two files. dirmap's other differentiator against both is the crawl mode that builds a dictionary from the target's own links via XPath, which neither of the flag-driven tools does in the same way.

Licence, releases and what maintenance looks like

dirmap is licensed under GPL-3.0. For anyone embedding it in another tool, that is a copyleft licence, and the practical question of whether your use counts as distribution is one for a lawyer, not for this article. The repository itself carries the LICENSE file at the top level.

On maintenance, the facts are narrow. The repository is not archived. The last push to master was on 2026-09-16, so commits are recent. The only release listed is v1.1, dated 2024-10-24. Those two facts pull in different directions, and the README explains part of why: the TODO list still has open items, and several configuration keys are marked not implemented. A recent push does not by itself mean the unimplemented features have landed, and the README as given still describes them as missing.

The upgrade cost is low in one sense and uncertain in another. There is no documented migration path between versions, and the README's own update mechanism, conf.update, is marked not implemented, so upgrading means pulling from git. Because the configuration is a plain ini-style file with named keys, an upgrade that adds a key will not break existing ones, but an upgrade that renames a key will fail silently: the README does not describe validation of unknown or missing keys, so a typo or a stale key name is not reported.

Editorial conclusion

Adopt dirmap if you want a Python scanner whose scan modes, status-code filters and request behaviour live in one dirmap.conf file you can diff and version alongside a target list, and if you are comfortable editing that file because the README states detailed configuration is not available through command-line arguments. Do not adopt it if you need a maintained release cadence to point at: the only release listed is v1.1 from 2024-10-24, even though the last push to master was on 2026-09-16. Before relying on it, verify three things yourself: that python3 -m pip install -r requirement.txt resolves on your Python version, that conf.dict_mode and the dictionary path in dirmap.conf point at a wordlist you actually have, and that the output folder writes one deduplicated file per target domain as the README claims.

Frequently asked questions

What is the DIR command used for?

In the context of dirmap, the tool is run as python3 dirmap.py with flags such as -i for the target and -lcf to load dirmap.conf; the README does not describe a DIR command of its own.

What is a DIR list?

dirmap does not use the term DIR list. Its equivalent is the dictionary, loaded from conf.dict_mode_load_single_dict in dict mode or from a directory of wordlists set by conf.dict_mode_load_mult_dict.

What is the purpose of a directory?

dirmap scans web directories and files: it requests paths built from a wordlist and records the responses whose status codes appear in conf.response_status_code, which defaults to [200].

What is an example of a directory?

The README's example target is https://target.com, scanned with python3 dirmap.py -i https://target.com -lcf. Recursion into discovered paths is controlled by conf.recursive_scan and conf.recursive_status_code.

Official sources

  1. H4ckForJob/dirmap on GitHub
  2. Issues
  3. License: GPL-3.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/h4ckforjob-dirmap.svg)](https://hysenlabs.com/projects/h4ckforjob-dirmap)