# gitGraber: watching GitHub search results for leaked credentials

> A small Python script that polls GitHub code search with your own keywords and token patterns, then pushes anything that looks like a credential into Slack, Discord or Telegram on a cron job.

**hisxo/gitGraber** — gitGraber: monitor GitHub to search and find sensitive data in real time for different online services such as: Google, Amazon, Paypal, Github, Mailgun, Facebook, Twitter, Heroku, Stripe...

- Repository: https://github.com/hisxo/gitGraber
- Stars: 2,438 · Forks: 370
- Language: Python
- License: GPL-3.0
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/hisxo-gitgraber

## It monitors search results, not repository history

The README is upfront about what the tool is not. It says plainly that gitGraber is not designed to check the history of repositories, and that many tools already do that well. What it does instead is monitor and parse the most recently indexed files on GitHub, which is a different and arguably more useful angle for leak hunting.

The distinction matters because the two approaches answer different questions. A history scanner asks what a specific repository once contained. gitGraber asks what is appearing across GitHub right now that matches your vocabulary, whether or not you already know the repository exists.

The README also states the reasoning behind that design. Its claim is that leaks come not only from organizations themselves but from service providers and employees who have no profile indicating where they work, which is exactly the population a keyword search reaches and a known-repository scan misses. GitHub search is therefore the primary index rather than an implementation detail, and the tool is built around constructing a good query rather than around crawling.

## A Python script with four dependencies

The repository is seven entries: a `COPYING` file, the README, `config.py`, `gitGraber.py`, `requirements.txt`, `tokens.py` and a `wordlists/` directory. That is the entire codebase, and the README recommends installing dependencies with pip against the requirements file, whose contents are short enough to read in full:

```bash
requests==2.32.0
argcomplete==3.2.1
python_crontab==2.3.9
termcolor==1.1.0
```

Each dependency maps to a visible feature. `requests` is the HTTP layer. `argcomplete` gives shell completion for the command line arguments, which is a small quality signal for a tool with this many flags. `python_crontab` is how the monitoring mode installs itself as a scheduled job. `termcolor` colours the CLI output so a hit is visually distinct from routine status lines.

The project is GPL-3.0 licensed, written in Python 3, and its topics are bugbounty, leaks, monitor, osint, realtime, redteam, security-automation and security-tools. There are no releases published, no tests directory and no build pipeline, which is consistent with a single-file script distributed by clone.

## Queries are keywords times a search term

The command line takes a keyword file and a query, and combines each word in the file with the query. The README's example searches for a specific word in combination with every word of `keywordsfile.txt` and sends output to Slack:

```bash
python3 gitGraber.py -k keywordsfile.txt -q YOURWORD -s
```

The help output lists the flags the script accepts, including a keyword file flag, a query flag, a monitor flag that creates a cron job running every thirty minutes, notification flags for Discord, Slack and Telegram, a wordlist flag, and a limit flag that restricts results to commits less than a given number of days old.

Two details in the README are worth repeating because they are the ones people get wrong. Domain names have to be wrapped in double quotes when used as the query, because otherwise GitHub search treats the dots as syntax. And the day limit exists because an unbounded search over a broad keyword will otherwise return matches from years ago, which is noise when you are trying to catch a fresh leak.

## Notification channels and what they cost you

Notifications go to Slack, Discord or Telegram, one channel at a time, chosen by flag. A hit also prints to the command line, so the channels are additive rather than exclusive.

Configuration lives in `config.py` and the README lists what you have to fill in: your own GitHub personal access tokens, a Discord webhook URL, a Slack webhook URL and a Telegram config holding a bot token and a chat id. The table underneath links to the official pages for creating each of those credentials.

The GitHub tokens deserve the extra attention, and the README has a whole FAQ entry about them. When GitHub detects a large number of requests from your token, the tool reports abuse detection reached and moves on to another token defined in the config file. The README is explicit that this is a temporary limit and that you do not need to create a new token, which tells you token rotation is built in rather than being something the operator manages. It also explains why you might see only a GitHub query and a status code in the output: a status of 200 means the request succeeded, and nothing else appeared because no token pattern matched your keyword.

## Detection quality is the regex list, and the list is old

The patterns live in `tokens.py` and the README is candid about their quality. It says the regular expressions are meant to be as accurate as possible, acknowledges there will be false positives, invites contributions to add patterns, and states a deliberate preference: the project would rather miss a standard API key than send a notification about one.

The blacklist mechanism is also documented. To ignore a specific pattern for a given token, you edit `tokens.py` and add the pattern as a list argument when that token is initialised, so the token's own constructor call gains an extra list. That is the extension point for anyone who finds a noisy provider.

The list itself is the problem. The README claims 31 supported tokens and dates that claim to 12 September 2019, and the names it then prints are shorter than 31: AWS, Facebook, a GitHub client secret, several Google variants, Heroku, JSON Web Tokens, Mailchimp, Mailgun, PayPal, and a run of private key types covering SSH, RSA, DSA, Elliptic Curve, PGP and OpenSSH. A count that does not match the list it introduces is a small thing on its own, but combined with a 2019 date it tells you to audit the file before relying on it. The last recorded push to the repository is 2026-03-26.

## Cron monitoring, dynamic wordlists and repeat suppression

The monitoring mode is the feature that makes the tool more than a one-shot search. With the monitor flag, gitGraber creates a cron job based on your query that runs every thirty minutes, and the README states the effect plainly: it searches for secrets on that query every half hour and sends a Slack notification whenever there are hits.

The wordlist flag does something different and is easy to miss. Rather than searching, it builds a wordlist that fills dynamically with filenames discovered on GitHub, which the README suggests feeding to your favourite fuzzing tool. That turns the tool into a filename discovery step in a larger pipeline rather than a scanner in isolation.

Deduplication is handled by a file. The README says gitGraber stores all repository URLs in a file named `rawGitUrls.txt`, so a repository already scanned and found to contain a key will not notify you a second time. The consequence is that this file is state: keep it if you want history, and expect a fresh run against a clean checkout to re-report what you have already seen.

## Conclusion

gitGraber does one narrow thing well: it turns GitHub search into a repeating query over your own keywords and your own regex list, and it tells you when something turns up. The design choice that separates it from a history scanner is that it watches newly indexed files rather than repository commits, which suits leak hunting where the interesting artifact is a config file someone pushed last week. The limits are equally clear. Detection quality is entirely down to the patterns in `tokens.py`, the README notes the list has not been updated since September 2019, and the last recorded push to the repository is 2026-03-26, so the GitHub search API it depends on may have moved underneath it. Install the four dependencies from `requirements.txt`, put your tokens in `config.py`, start with one keyword and one notification channel, and verify the search syntax still returns results before trusting a quiet run.

## FAQ

### How can I scan a Git repository for secrets?

gitGraber does not scan repository history, by its own account, since other tools already do that well. It monitors the most recently indexed files on GitHub by searching for your own keywords combined with a query term, then applying the token patterns in `tokens.py` to whatever comes back.

### Which API keys and services does gitGraber detect?

The README claims support for 31 tokens with patterns held in `tokens.py`, dated 12 September 2019. The names it lists include AWS, Facebook, GitHub client secret, several Google variants, Heroku, JSON Web Tokens, Mailchimp, Mailgun, PayPal, and private keys in SSH, RSA, DSA, EC, PGP and OpenSSH formats.

### How do I set up notifications for gitGraber?

You edit `config.py` to add your GitHub personal access tokens plus whichever of a Discord webhook URL, a Slack webhook URL or a Telegram bot token and chat id you want, then pass the matching flag such as `-s` for Slack when you run the script. When GitHub rate limits a token the tool rotates to another one from the config file.

### What does the monitor flag do in gitGraber?

It creates a cron job based on your query that runs every thirty minutes, using the python_crontab dependency, so the same search repeats and a notification is sent whenever there are hits. Results are deduplicated through a file named `rawGitUrls.txt` so an already reported repository does not notify twice.

### How do I filter out a false positive in gitGraber?

The README describes editing `tokens.py` and adding the pattern as a list argument when that token is initialised, which becomes an extra list in the token's constructor call. That blacklist is the intended way to silence a provider whose keys match too broadly.

## Sources

- [hisxo/gitGraber on GitHub](https://github.com/hisxo/gitGraber)
- [Issues](https://github.com/hisxo/gitGraber/issues)
- [License: GPL-3.0](https://github.com/hisxo/gitGraber/blob/master/LICENSE)
- [README](https://github.com/hisxo/gitGraber/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/hisxo-gitgraber
