internetwache/GitTools: three shell and Python tools for exposed .git directories
A repository with 3 tools for pwn'ing websites with .git repositories available
At a glance
- What is it?
- GitTools bundles a finder, a dumper and an extractor for websites that leave their .git directory public. The dumper is the practical core; the finder is a mass scanner and the extractor is a repair step for incomplete downloads.
- Who is it for?
- Adopt GitTools when you need to pull a publicly exposed .git directory from a server that returns 403 on directory listing, or when you already hold a partial repository and need the extractor to rebuild commit contents. Do not adopt the Finder as a general-purpose scanner for a large host list unless you have written permission for every target, and do not expect the Dumper to recover packed repositories reliably, since the README states it may fail on pack-files.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 77 days ago.
- What is it written in?
- Mainly Shell, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What GitTools is for, and who should pick it up
GitTools is a collection of three small scripts, written in Python and bash, that deal with one situation: a web server that serves its own .git directory to anyone who asks. The README frames the whole repository as tooling for what it calls "Git research", and points to a write-up about an analysis of Alexa's top one million sites. The three pieces split the work by stage. Finder locates hosts whose .git/HEAD is reachable. Dumper pulls down as much of that repository as the server will hand over, including servers that do not enable directory listing. Extractor repairs an incomplete repository by walking its commit objects and restoring their contents.
The intended user is someone doing security assessment with authorization over the targets, or a developer recovering their own accidentally published repository. The Finder's own help text is explicit that usage "might be illegal in certain circumstances" and that it is meant for educational purposes. That line is not boilerplate; it is the boundary of the project. If you cannot point to permission for a host, the Finder has nothing to offer you. The Dumper and Extractor are far less exposed, because they act on a URL or a directory you already chose.
How the three scripts divide the work
The mechanism differs per tool, and the differences matter when you decide which one to run.
Finder is a Python script that takes a list of targets, one per line, and checks a single condition: whether the .git/HEAD file contains the string refs/heads. That is a narrow test. A server that returns a custom 200 page for every path will not pass it, and a server that blocks .git/HEAD but leaks other files under .git will not be reported. The script runs with a thread count you supply and prints matches to stdout as lines of the form [*] Found: DOMAIN.
Dumper is a bash script that takes a base URL and a destination directory. Its distinguishing feature is that it does not need directory listing: it requests the files it expects to exist inside a .git directory rather than enumerating them. The README states plainly that it has "no 100% guaranty to completely recover the .git repository" and that it may fail when the repository has been compressed into pack-files. That is the honest description of the design. Loose object files are recoverable one by one; a packed repository is a single opaque blob and the dumper cannot reconstruct it from the outside.
Extractor is a bash script that takes a repository directory and a destination directory. It iterates through all commit objects, tries to restore the contents of each commit, and writes the results out. The README notes that commits are not sorted by date, which follows from the fact that it is enumerating objects rather than walking a graph from a known HEAD. The README presents it as a companion to the Dumper for the case where the downloaded repository is incomplete.
Installing GitTools and dumping a first repository
There is no package to install. The tools live in the repository as scripts, so you clone the repository and run them from the checkout. The README lists the runtime requirements: git, Python 3 or newer, curl, bash, sed, binutils (for strings) and file. On a Debian or Ubuntu machine, the binutils and file packages are the two that are often missing from a minimal image.
After cloning, the Dumper is invoked with the target's .git URL and a destination directory. The help output shows the optional --git-dir flag for targets that renamed the directory:
./gitdumper.sh http://target.tld/.git/ dest-dirThe destination directory is created by the script and receives whatever it manages to fetch. The same help output documents the --git-dir flag, which changes the git folder name from the default .git.
Once the download finishes, the Extractor turns the fetched objects into readable commit contents. The README gives this exact invocation:
./extractor.sh /tmp/mygitrepo /tmp/mygitrepodumpHere /tmp/mygitrepo is the directory that contains the .git directory, and /tmp/mygitrepodump is where the restored contents land. The README says this combination is the intended use when the downloaded repository is incomplete. Expect the output to be unsorted by date, because that is how the script enumerates commits.
The Finder is the third script. The README's own example of building an input file from Alexa data uses wget, unzip and sed to strip the rank prefix from each CSV line, so the file ends up as one bare hostname per line. The script's help output documents the -i, -o and -t options:
./gitfinder.py -i INPUTFILE -o OUTPUTFILE -t THREADSThe -i flag names the input file with one target per line, -o names the output file, and -t sets the thread count. Matches appear on stdout as [*] Found: DOMAIN.
Where GitTools stops working
The pack-file limitation is the one the README admits to, and it is the one that bites most often in practice. Servers that run git gc, or that were deployed from a cloned repository rather than a fresh init, will hold their history in pack-files. The Dumper can still fetch the loose files it knows about, but the README states it may fail to recover the repository in that case. There is no fallback described for unpacking a pack-file from the outside, and the Extractor works on commit objects, not on packs.
The Finder's detection rule is a second boundary. Checking only .git/HEAD for refs/heads produces false negatives on servers that serve .git/HEAD with a different content type or behind a rewrite rule, and it produces no information about whether the rest of the directory is reachable. A host that passes the Finder test can still yield almost nothing to the Dumper.
The Extractor is also the wrong tool for a healthy repository. If you have a working clone, git log and git checkout already give you the history in order. The Extractor exists for the case where the object database is incomplete and the normal git commands refuse to run. Running it against a complete repository adds nothing.
Finally, the whole set is the wrong choice for scanning at scale without authorization. The README's Alexa example is a historical research exercise, and the help text's warning about legality is the project telling you where it draws the line.
GitTools compared with GitVersion and other tools that share the name
Search traffic for this project collides with a different family of software. GitVersion, distributed under the GitTools organization on GitHub, is a semantic versioning tool: it reads your repository history and computes a version number for your build, typically wired into Azure DevOps pipelines or GitHub Actions. That is a completely different job from what internetwache/GitTools does. If you arrived here looking for version calculation in CI, this repository is not it, and the two share only a name.
Within the same problem space, the closest alternative to the Dumper is git itself. A server with directory listing enabled can be mirrored with a plain recursive fetch, and git clone works against a server that exposes a usable smart or dumb HTTP endpoint. The Dumper's reason to exist is the middle case: directory listing is off, the smart protocol is not available, but individual files under .git are still served. That is a narrow window, and it is exactly the window the README describes.
For the Extractor, the alternative is git fsck and git cat-file against the partial object database. Those commands tell you what is missing but leave you to reconstruct the contents. The Extractor automates the reconstruction step and accepts that the result is unsorted.
Proxy setup, licence and what maintenance looks like
Both the Python and the shell side reach the network through tools that honor standard proxy environment variables. The README says urllib and curl should pick up HTTP_PROXY and HTTPS_PROXY, and gives the bash and Windows CMD forms. In bash:
export HTTP_PROXY=http://proxy_url:proxy_port
export HTTPS_PROXY=http://proxy_url:proxy_portBasic authentication is supported by embedding credentials in the proxy URL, in the form http://username:password@proxy_url:proxy_port. This matters for the Finder, which makes many requests and will otherwise fail silently on a network that requires a proxy.
The licence is MIT for all three tools, with the text in LICENSE.md. MIT is permissive: it allows reuse and modification with the copyright notice and permission notice retained. It offers no warranty, and it says nothing about the legality of what you point the tools at. That question sits with you and with the jurisdiction you are working in, not with the licence.
On maintenance, the repository is not archived, and the most recent push recorded for the default branch is 2026-07-15. The only tagged release is v0.0.1 from 2021-05-05, described as the initial release. The scripts are small and depend on stable interfaces (git object layout, curl, the .git directory structure), so the low release cadence is not by itself a problem, but the version number tells you the project has never declared a stable interface.
Editorial conclusion
Adopt GitTools when you need to pull a publicly exposed .git directory from a server that returns 403 on directory listing, or when you already hold a partial repository and need the extractor to rebuild commit contents. Do not adopt the Finder as a general-purpose scanner for a large host list unless you have written permission for every target, and do not expect the Dumper to recover packed repositories reliably, since the README states it may fail on pack-files. Before relying on it, verify three things on a target you control: that gitdumper.sh writes into the destination directory you pass, that extractor.sh produces commit contents in /tmp/mygitrepodump, and that your proxy environment variables are picked up by both urllib and curl.
Frequently asked questions
What does internetwache/GitTools actually download?
The Dumper fetches as much as it can from a target's .git directory when directory listing is disabled. The README warns it has no guarantee of recovering the whole repository and may fail when the history is stored in pack-files.
How do I run the GitTools extractor on a downloaded repository?
The README gives the invocation ./extractor.sh /tmp/mygitrepo /tmp/mygitrepodump, where the first path contains a .git directory and the second is the destination. Commits are restored but are not sorted by date.
Is internetwache/GitTools the same as Gittools GitVersion?
No. GitVersion is a semantic versioning tool used in build pipelines. internetwache/GitTools is a set of three scripts for finding, downloading and extracting exposed .git repositories.
What are the requirements for running GitTools?
The README lists git, Python 3 or newer, curl, bash, sed, binutils for the strings utility, and file. There is no package to install; the scripts are run from a clone of the repository.
Can GitTools use a proxy?
Yes. The README states that urllib and curl support proxy configuration through the HTTP_PROXY and HTTPS_PROXY environment variables, and that basic authentication works by embedding credentials in the proxy URL.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/internetwache-gittools)