Model or dataset
apurvsinghgautam/robin avatar
apurvsinghgautam/robin

Robin: an AI-assisted dark web OSINT tool that filters search results before you read them

AI-Powered Dark Web OSINT Tool

7,386 stars1,382 forksPythonMIT

At a glance

What is it?
Robin pairs a Tor-backed search and scrape pipeline with an LLM filtering stage, so an investigation returns a shortlist and a summary instead of a wall of onion links. The Docker image is the recommended install; the honest-results behaviour is the part worth judging.
Who is it for?
Adopt Robin if you already have a lawful investigative remit, a provider API key or a local Ollama instance, and you want the filtering and summary step handled for you. Do not adopt it if you cannot send your queries to a third-party model provider, or if you need a documented rollback path: the README covers installation and troubleshooting but says nothing about undoing a run or auditing what was sent to the provider.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Robin targets: dark web search results are cheap, triage is not

Dark web search engines return links, not answers. A query about a leaked dataset or a stolen credential dump produces a page of onion addresses whose titles are often misleading, duplicated across mirrors, or dead by the time you click. The expensive part of an OSINT investigation is not finding candidate pages. It is reading them, deciding which ones are relevant, and writing down what you found in a form someone else can check.

Robin is built for that second half. The README describes it as "an AI-powered tool for conducting dark web OSINT investigations" that uses LLMs to refine queries, filter search results from dark web search engines, and produce an investigation summary. The target user is someone with a lawful investigative remit: a threat intelligence analyst, a security researcher, a journalist working a specific story. It is not a general web search replacement, and the disclaimer in the README is explicit that accessing certain dark web content may be illegal depending on jurisdiction and that the author is not responsible for misuse.

One design decision stands out against the usual pattern for tools in this category. The README lists an "Honest Results" feature: when nothing relevant is found, Robin says so instead of summarizing whatever it happened to scrape. That is a small sentence with large consequences for how much you can trust the output.

How the search, scrape and LLM stages fit together

The repository layout matches the three-stage pipeline the README describes. search.py handles querying the dark web search engines, scrape.py fetches page content, and llm.py plus llm_utils.py wrap the model providers. config.py holds defaults, model_registry.py and models.json back the model discovery, and ui.py is the Streamlit front end. The README calls this a "Modular Architecture" with clean separation between search, scrape, and LLM workflows, and the file names support that reading.

Data flows one direction. A query goes out through Tor to the search engines; the returned links are filtered by the model down to the number of results you set in the sidebar; the surviving pages are scraped, with a cap on how many pages and how much of each page the model reads; then the model writes the summary. The sidebar exposes all three dials and, per the README, shows the token cost before you run. That pre-run estimate is the most practically useful control in the interface, because scraping depth and model choice are the two things that decide whether a run costs cents or dollars.

Model selection is dynamic rather than hardcoded. The README states that models are discovered from each provider at startup, so new releases appear on their own and retired ones disappear, with no hardcoded list to go stale. Supported providers are OpenAI, Claude, Gemini, Mistral, OpenRouter, Ollama, and any OpenAI-compatible API such as LM Studio, llama.cpp or Groq. Two features sit on top of the pipeline: conversational follow-ups, answered from that investigation's own data without re-running the search, and one-click pivots that turn a suggested follow-up query into a fresh investigation.

Installing Robin with Docker and running a first investigation

The README recommends Docker and states that the tool needs Tor to do the searches. On Linux or Windows under WSL that is `apt install tor`; on Mac it is `brew install tor`. Confirm Tor is running in the background before you start.

You need at least one provider key. Copy `.env.example` to `.env` and fill in whichever of the supported keys you have. One key is enough, because Robin lists models for whichever providers it finds.

bash
OPENAI_API_KEY=
ANTHROPIC_API_KEY=
GOOGLE_API_KEY=
MISTRAL_API_KEY=
OPENROUTER_API_KEY=

Pull the image, then run it with the `.env` file mounted. The `--add-host` flag is required for the Ollama path described below.

bash
docker pull apurvsg/robin:latest
bash
docker run --rm \
   -v "$(pwd)/.env:/app/.env" \
   --add-host=host.docker.internal:host-gateway \
   -p 8501:8501 \
   apurvsg/robin:latest

Open `http://localhost:8501`. Choose a provider and model, set how many results to filter, how many pages to scrape, and how much of each page the model reads, check the token estimate, and run the investigation. Saved investigations land in an `investigations/` folder; the README shows a second `docker run` with `-v "$(pwd)/investigations:/app/investigations"` so they survive a restart, and they can be reloaded from the Past Investigations panel in the sidebar.

The Python route is for development. With Python 3.10+ and Tor installed, `pip install -r requirements.txt` followed by `streamlit run ui.py` gives the same interface on the same port.

Ollama and custom providers: where the setup actually bites

Running a local model is the option that keeps queries off a third-party API, and it is also the option with the most moving parts. The README states that nothing goes in `.env` for Ollama, because Robin defaults to `http://host.docker.internal:11434`, which is what the Docker install needs. Two things remain the user's responsibility.

First, the container needs `--add-host=host.docker.internal:host-gateway`, which the recommended command already includes. Second, Ollama binds to `127.0.0.1` by default and a container cannot reach that, so Ollama has to listen on all interfaces. Started by hand, that is `OLLAMA_HOST=0.0.0.0 ollama serve &`. Under systemd, the README's route is `sudo systemctl edit ollama.service`, adding `[Service]` and `Environment="OLLAMA_HOST=0.0.0.0"`, then `sudo systemctl daemon-reload && sudo systemctl restart ollama`.

If you run Robin directly with Python instead of Docker, the default is wrong for your setup and you override it with `OLLAMA_BASE_URL=http://127.0.0.1:11434`. The README points at TROUBLESHOOTING.md if the model still does not appear. That document is also the first stop for an empty model dropdown, Tor `resolve failed` errors, and 401 responses.

For any other OpenAI-compatible provider, the README says to use the Custom API Provider expander in the sidebar, entering a base URL, an optional API key, and optionally a model name if the provider does not expose `/v1/models` for auto-discovery. `.env.example` lists the corresponding overrides, including `CUSTOM_API_BASE_URL`, `CUSTOM_API_KEY` and `CUSTOM_API_MODEL`, plus `OLLAMA_NUM_CTX` for a RAM-constrained machine and `OPENROUTER_BASE_URL` for a custom OpenRouter gateway. The pattern is consistent: `.env` for defaults you want everywhere, the sidebar for one-off endpoints.

Where Robin is the wrong tool

The filtering step is the product, and it is also the failure mode. An LLM decides which search results are relevant before you ever see them. If the model's judgement is wrong in the direction of exclusion, you will not know what you missed, because the discarded links are not the ones that reach the summary. The tunable depth settings are a mitigation, not a fix: raising the number of results to filter and the pages to scrape widens the net at a direct token cost, and the README gives no way to inspect the rejected set.

The honest-results behaviour cuts the other way and is worth weighing carefully. A tool that reports "nothing relevant found" is more trustworthy than one that always produces a summary, but it also means a null result is a real output you have to interpret. It could mean the material is not there, or that the query was wrong, or that the model filtered out the page that mattered. Robin does not resolve that ambiguity for you.

There is also a data-handling boundary. The README warns that Robin uses third-party APIs including LLMs and advises caution when sending potentially sensitive queries, with a review of the terms of service for any provider you use. If your investigation cannot leave your infrastructure, the only configurations that respect that are Ollama or another self-hosted OpenAI-compatible endpoint. Anything else sends your query text to a vendor.

Finally, the README does not document rollback. There is no described way to reverse a run or to audit exactly what was transmitted to a provider, which matters if your work has a retention or disclosure obligation.

How Robin differs from running a search engine and a summarizer yourself

The obvious alternative is to do the three stages by hand: query a dark web search engine, open the results in Tor Browser, and paste the pages you care about into a general-purpose LLM chat. That approach gives you full control over what the model sees and lets you read every candidate page. Its cost is time and consistency. You are the filter, and two analysts doing the same query will produce different shortlists.

Robin's difference is that the filter, the scraper and the summarizer share one configuration and one interface, with the depth settings and the token estimate visible before the run. The follow-up questions are answered from the stored investigation rather than from a fresh search, which keeps a conversation grounded in what was actually collected. The pivots turn a finding into a new query in one click. Neither of those is a capability a chat window has, because a chat window has no investigation object to be grounded in.

What you give up is transparency. A manual workflow lets you read the pages the model never saw. Robin's design assumes you would rather not. For a narrow, well-specified question that trade is reasonable; for exploratory work where you do not yet know what you are looking for, it is the wrong shape.

Maintenance, releases and the MIT licence

The repository is not archived and the last push was on 2026-09-11, the same day as the v2.9 release. The release history shows a steady cadence: v2.7 on 2026-06-02, v2.8 on 2026-07-15, v2.9 on 2026-09-11. Nothing in the README describes a support policy, a deprecation window or a compatibility guarantee between versions, so treat the release notes as the only record of what changed.

Upgrade cost is concentrated in the provider layer. Because models are discovered at startup rather than pinned in a list, a provider retiring a model changes your dropdown without a Robin release. The README frames that as an advantage, and for staying current it is. It also means the thing that breaks is outside the project's control, and the fix is usually to pick a different model rather than to change code. Pinning a specific model through the Custom API Provider fields is the way to avoid that class of surprise.

The project is MIT licensed. That is permissive in the usual sense: it allows commercial and closed-source use, modification and redistribution provided the copyright notice and permission notice are kept. It says nothing about the legality of the material you point the tool at, and the README's disclaimer separates those two questions deliberately. The licence covers the code, not your investigation.

Editorial conclusion

Adopt Robin if you already have a lawful investigative remit, a provider API key or a local Ollama instance, and you want the filtering and summary step handled for you. Do not adopt it if you cannot send your queries to a third-party model provider, or if you need a documented rollback path: the README covers installation and troubleshooting but says nothing about undoing a run or auditing what was sent to the provider. Verify first that Tor is running, that your chosen provider appears in the model dropdown, and that the token estimate shown in the sidebar matches what you are willing to spend before you press run.

Frequently asked questions

How do I install Robin AI?

The README recommends Docker: install Tor, copy .env.example to .env and add at least one provider API key, then pull apurvsg/robin:latest and run it with the .env file mounted and port 8501 published. A Python route exists for development, using Python 3.10+ and `streamlit run ui.py`.

Does Robin need Tor to work?

Yes. The README states that the tool needs Tor to do the searches, installed with `apt install tor` on Linux or Windows under WSL, or `brew install tor` on Mac, and that you should confirm Tor is running in the background. The Dockerfile installs Tor as part of the image build.

Which LLM providers does Robin support?

OpenAI, Claude, Gemini, Mistral, OpenRouter, Ollama, and any OpenAI-compatible API such as LM Studio, llama.cpp or Groq. The README states that models are discovered from each provider at startup, so the list of available models is not hardcoded.

Can I run Robin with a local model instead of a cloud API?

Yes, through Ollama. The README notes that nothing goes in .env for Ollama because Robin defaults to http://host.docker.internal:11434 for the Docker install, but you must run the container with --add-host=host.docker.internal:host-gateway and make Ollama listen on all interfaces rather than 127.0.0.1.

What should I do if the model dropdown is empty or Tor reports resolve failed?

The README directs you to TROUBLESHOOTING.md for an empty model dropdown, Ollama not appearing, Tor resolve failed errors, 401 responses and no-results situations. It asks that you read that file before opening an issue.

Official sources

  1. apurvsinghgautam/robin on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/apurvsinghgautam-robin.svg)](https://hysenlabs.com/projects/apurvsinghgautam-robin)