# Maigret: a default run checks 500 of the 5,900 sites it supports

> Maigret builds a dossier on a person from a username alone, with no API keys, by checking sites and extracting profile data. The default run touches the 500 highest ranked sites out of 5,900 supported, the site database is refreshed from GitHub once a day, and the Docker image binds its web interface to every network interface.

**soxoj/maigret** — 🕵️‍♂️ Collect a dossier on a person by username from 3000+ sites

- Repository: https://github.com/soxoj/maigret
- Website: https://maigret.app
- Stars: 38,189 · Forks: 3,009
- Language: Python
- License: MIT
- Published: 2026-08-17 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/soxoj-maigret

## A default run checks 500 sites, not the 5,900 it supports

The feature list is unambiguous about this. Maigret supports 5,900 sites, with the full list kept in sites.md, and a default run checks the 500 highest ranked sites by traffic. Two flags change that: -a scans everything, and --tags narrows the set by category or country.

The entry point is two commands and a version floor:

```bash
pip install maigret
maigret YOUR_USERNAME
```

The consequence is that the default output and a complete output are different documents, and nothing in the default run marks the difference. A username with no results in the top 500 and a username with no results anywhere look the same unless you know that five fifths of the database was never queried. If the answer matters, -a is the run that supports a conclusion, and it is also the run whose duration nobody in the repository has documented. The site's own description is out of step too, still saying 3000+ sites while the README says 5,900.

## The site database is fetched from GitHub on every run

Maigret keeps its list of sites fresh by downloading an updated database from GitHub, once per 24 hours, and it falls back to the built in database when it cannot reach the network. That is the documented behaviour, and it is the reason two people running the tool on the same day can get different site lists.

The consequence is threefold. A first run of the day is slower than a later one because it pays for the fetch, a run on an isolated or air gapped machine quietly uses whatever database was built into the package, and the version of the data behind a result is not printed in the report. For a workflow where a finding matters, the age of the site list is part of the evidence, and the tool does not record it. Sites also come and go, so a negative result means the site was checked and found nothing, not that the site still exists.

## The README opens with four proxy vendors and their discount codes

Above the feature list sit four sponsor blocks, and three of them sell proxies. Noimosiny is presented as an OSINT platform for investigators doing reverse email, phone, and username search. RapidProxy advertises residential proxies for Twitter scraping, Selenium automation, and web data extraction, at plans from $0.65/GB with the code RAPID10. MangoProxy sells residential, ISP, mobile, and datacenter proxies with the code SOXOJ. Thordata advertises residential IPs for OSINT research with the code SOXOJ10.

The consequence is that the page selling you a way around blocks is the same page that concedes the blocks exist. The feature list says outright that Maigret detects and partially bypasses blocks, censorship, and CAPTCHA, and the word partially is doing real work there. That means coverage depends on your network egress and on how a given site responds today, which is why the vendor list is not incidental to the tool. A run from a residential IP and a run from a datacentre IP are not the same run, and the report does not record which one produced it.

## The image binds the web interface to 0.0.0.0 and defaults to the CLI

The Dockerfile has three stages, and the ordering matters. The base stage is python:3.11-slim with the build dependencies for lxml and cairo, and the install step sets YARL_NO_EXTENSIONS=1 before pip install, which forces the pure Python path for aiohttp. Then there is a web stage that installs the pdf extra, sets PORT to 5000, exposes it, and runs maigret --web on that port. Finally there is a cli stage, annotated as the last stage and therefore the target of a plain docker build.

The security detail is in the base stage, which sets FLASK_HOST to 0.0.0.0 and carries a comment saying that for production use you should set FLASK_HOST to a specific IP address.

The consequence is that docker build . produces the command line tool, not the web interface, so reaching the browser version means naming the target explicitly. And the image as written listens on every interface, so publishing that port without overriding FLASK_HOST puts a username search interface on whatever network the container is attached to, which is exactly what the comment warns against in the one line most readers skip.

## The only timing target in the repository is import time

The Makefile has a speed target, and it measures startup rather than searching:

```bash
time python3 -m maigret --version
python3 -c "import timeit; t = timeit.Timer('import maigret'); print(t.timeit(number = 1000000))"
python3 -X importtime -c "import maigret" 2> maigret-import.log
```

A million import iterations, the importtime log, and then python3 -m tuna to read that log. Nothing in the repository times a 500 site scan or a full -a scan, so there is no figure to plan a batch around.

The consequence is that the project optimises the thing a user notices once and leaves the thing a user waits for unmeasured. Import time matters for CLI responsiveness, and a million repetitions is a deliberate way to measure a fraction of a second. Runtime against live sites is dominated by the sites rather than by the code, which is a reasonable thing not to benchmark and a poor thing to leave undocumented, because a researcher deciding whether to run -a overnight has nothing to go on. The other Makefile targets are ordinary: pytest under coverage, flake8, mypy, black, pip3 install ., and a clean that removes reports, htmcov, and dist.

## The Windows build is a separate artifact from the pip package

Python 3.10 is the floor everywhere it is stated: the quick start asks for 3.10 or higher, the pyproject constraint is ^3.10, and the classifiers cover 3.10 through 3.14. The container is built on 3.11.

Windows users get something different. The install path is a standalone executable downloaded from the releases page, launched either by double clicking it, which asks for a username and waits at the end so the report links stay on screen, or from a terminal:

```cmd
cd %USERPROFILE%\Downloads
maigret_standalone.exe USERNAME
maigret_standalone.exe USERNAME --html       :: also save an HTML report
maigret_standalone.exe --help                :: list all options
```

The repository carries Installer.bat and a pyinstaller directory, which is how that binary is produced. The releases list mixes versioned tags, v0.6.6 and v0.6.5, with a rolling nightly-main Development Windows Release that carries no version number.

The consequence is that the Windows path is a frozen binary while everyone else installs from PyPI, so the two can drift, and a nightly with no number attached tells you nothing about which code it holds. A Windows user who wants the pip behaviour has to install Python first, and a pip user has no reason to look at the releases page at all.

## The dossier needs no keys, but the AI summary needs an endpoint

The core promise is that no API keys are required. Maigret works by username alone, extracts what it can from profile pages and site APIs through the socid-extractor dependency, follows discovered usernames and other identifiers recursively, and can browse results as a graph in a web interface and download reports in several formats from one page. Public PDF and HTML reports plus a full console transcript of a recursive search are published in the repository as examples of the output.

The optional part is separate. The --ai mode turns raw findings into a short investigation summary using an OpenAI compatible API, and it is opt in.

The consequence is a clean split worth respecting. The scan itself is local apart from the site database fetch, but turning findings into a summary sends them to whatever endpoint you configured, so a user whose findings must not leave the machine should leave --ai off and read the report. The tool also reaches Tor and I2P sites and can check domains, and the recursive search means a single username can pull in accounts you did not ask about, which is worth knowing before a first run on a name you care about. The table of contents also lists a Commercial Use section, and the visible README does not summarise it.

## Conclusion

Maigret fits an investigator or researcher who already knows what they are looking for, has a username, and wants a broad sweep with a text or HTML report at the end. It does not fit anyone who needs a complete answer from a default run, because five fifths of the supported sites are skipped, and it does not fit a deployment where findings must not leave the machine, because the site database is fetched from GitHub on each run and the optional AI mode sends raw findings to an OpenAI compatible endpoint. Before you rely on it, run with -a at least once to see the real coverage, read the Commercial Use section if your work is paid, and override FLASK_HOST before you expose the containerised web interface to any network.

## FAQ

### how to install maigret

Install Python 3.10 or higher, then run pip install maigret and maigret YOUR_USERNAME. The project is published on PyPI as maigret, and alternative paths include a Windows standalone executable from the releases page, cloud shells, a community Telegram bot, and a Docker image.

### how to use maigret

Pass a username: maigret YOUR_USERNAME runs a default search over the 500 highest ranked of the 5,900 supported sites. Pass -a to scan everything, --tags to narrow by category or country, --ai for an OpenAI compatible summary, and --html to also save an HTML report.

### how to install maigret on kali linux

The README does not mention Kali Linux. Its documented paths are pip install maigret on Python 3.10 or higher, the Windows standalone executable, browser based cloud shells including Google Cloud Shell, repl.it and Google Colab, a community Telegram bot, and a Docker image built from the repository Dockerfile.

### how to use maigret osint

Maigret is built for this: it needs only a username, requires no API keys, extracts account owner information from profile pages and site APIs with socid-extractor, and performs recursive search using discovered usernames and other IDs. It also runs against Tor and I2P sites and can check domains.

## Sources

- [Official documentation](https://maigret.app)
- [Official README](https://github.com/soxoj/maigret#readme)
- [Project repository](https://github.com/soxoj/maigret)
- [Release notes](https://github.com/soxoj/maigret/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/soxoj-maigret
