# Automated-AI-Web-Researcher-Ollama: a local-LLM research loop that searches and scrapes on its own

> A Python program that pairs Ollama with DuckDuckGo search and page scraping, then keeps researching until you stop it. It is a research log generator, not a fact checker.

**TheBlewish/Automated-AI-Web-Researcher-Ollama** — A python program that turns an LLM, running on Ollama, into an automated researcher, which will with a single query determine focus areas to investigate, do websearches and scrape content from various relevant websites and do research for you all on its own! And more, not limited to but including saving the findings for you!

- Repository: https://github.com/TheBlewish/Automated-AI-Web-Researcher-Ollama
- Stars: 3,018 · Forks: 277
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/theblewish-automated-ai-web-researcher-ollama

## What Automated-AI-Web-Researcher-Ollama actually automates

A single question to a chat model returns one answer with no trail. This project targets the opposite problem: you want many pages read and a record of where the text came from. According to the README, you type a query, the LLM generates five prioritized research focus areas, and it then works through them, formulating search queries, running web searches, picking pages, scraping them, and writing everything into a research text file that includes source URLs. The README states it can perform over a hundred searches and content retrievals depending on your system and model. The intended user is someone who already runs Ollama locally and wants an unattended research pass rather than a chat session. The README also notes a successor project, Academic-AI-Literature-Reviewer-Ollama, which it says searches academic articles instead of the internet, so this repository is the general web variant rather than the newest line of work.

## The research loop, from query to saved text file

The repository layout shows the loop is split across modules: Web-LLM.py is the entry point, research_manager.py coordinates the work, web_scraper.py handles retrieval, llm_wrapper.py talks to the model, llm_response_parser.py and strategic_analysis_parser.py interpret model output, and Self_Improving_Search.py is a separate search component. The README describes the cycle explicitly: the model analyzes your query, produces five focus areas with priorities, and starts with the highest priority one. For each area it formulates targeted search queries, performs web searches, analyzes the results, selects pages, and scrapes and extracts content. Findings are appended to a research text file with links. After the five areas are done, the README says the model generates new focus areas from what it found and repeats, which is why a session can run for a long time. The self-improving search mechanism is listed as a feature, and Self_Improving_Search.py exists in the tree, but the README does not document the algorithm behind it. Treat that file as the place to read before assuming how query refinement works.

## Installing it and running a first research session

The README documents Linux and macOS for the main branch and points Windows users to the feature/windows-support branch. Start by cloning the repository and creating a virtual environment, then install the dependencies from requirements.txt, which includes duckduckgo-search, beautifulsoup4, trafilatura and requests.

```bash
git clone https://github.com/TheBlewish/Automated-AI-Web-Researcher-Ollama
cd Automated-AI-Web-Researcher-Ollama
python -m venv venv
source venv/bin/activate
pip install -r requirements.txt
```

Install Ollama separately, following the instructions at ollama.ai. The README recommends a model with a large context window, naming phi3:3.8b-mini-128k-instruct or phi3:14b-medium-128k-instruct. Then open llm_config.py and set model_name to the model you have pulled, and adjust n_ctx for the context size you want.

```python
LLM_CONFIG_OLLAMA = {
    "llm_type": "ollama",
    "base_url": "http://localhost:11434",
    "model_name": "custom-phi3-32k-Q4_K_M",
    "temperature": 0.7,
    "top_p": 0.9,
    "n_ctx": 55000,
    "stop": ["User:", "\n\n"]
}
```

With Ollama serving, start the researcher and submit a query with a leading @, followed by CTRL+D. The README's example is @What year is the global population projected to start declining?.

```bash
ollama serve
python Web-LLM.py
```

During a session, the README lists single-letter commands submitted with CTRL+D: s for status, f for the current focus, p to pause and get an assessment of whether the model can already answer your query (then c to continue or q to quit), and q to stop. After you quit, the model reviews the collected content and produces a summary, then enters conversation mode for questions about the findings. The research session text file lands in the program's directory.

## Where this design breaks down

The research file is the real output, and its quality depends on scrapers and on a model deciding relevance. The README does not document any verification step: nothing checks a scraped claim against a second source, and the summary is generated from whatever text was collected. If a page is a content farm or a stale copy, that text enters the file and the summary. The README does not document rate limiting or retry behaviour for searches and scrapes, so a long session depends on how DuckDuckGo and the target sites respond. Context is the other hard constraint. The example config sets n_ctx to 55000 and the README recommends 128k-context models, which means small-context models will either truncate or fail once enough content accumulates. There is also no documented resume: the README describes quitting and summarizing, not restarting a session from a saved file. This is the wrong tool when you need reproducible, citable evidence, or when you want a quick factual answer that a single search would settle.

## How it differs from a search API plus a summarizer

A common alternative is to call a search API yourself, fetch the top results, and pass them to a model in one prompt. That approach gives you control over which pages are fetched and a fixed cost per run. This project inverts the control: the model decides the focus areas, the queries, and which results to open, and it keeps generating new focus areas as it learns, so the number of fetches is not fixed in advance. The trade-off is predictability. With a hand-written pipeline you know exactly what entered the prompt; here the research text file is the only record, and the README says it contains all retrieved content, source URLs, focus areas investigated, and the generated summary. If you need a deterministic audit trail, the manual pipeline is easier to defend. If you want breadth and are willing to review the log, the autonomous loop covers more ground per run.

## Maintenance, licence and the cost of upgrading

The repository is not archived, and the last push was on 2026-09-02, so it is recent enough to read as a live project rather than an abandoned one. There are no retrieved releases, so expect to track the main branch rather than pin a version. The licence is MIT, which permits commercial use and modification provided the copyright notice and permission notice are included; this is a description of the licence text, not legal advice, and you should read LICENSE in the repository before relying on it. The upgrade cost sits in two places: llm_config.py holds model name, base_url, temperature, top_p, n_ctx and stop sequences, so a model swap means editing that block and re-testing the loop, and requirements.txt pins no versions, which means a fresh install today can pull newer libraries than the code was written against. There is no documented migration guide for either.

## Conclusion

Adopt it if you already run Ollama locally, want an unattended research log with source URLs, and are willing to read the scraped pages yourself before trusting the summary. Skip it if you need Windows support from the main branch (the README points Windows users to the feature/windows-support branch), if you need citations you can verify programmatically, or if you cannot run a model with a large context window. Before committing, clone the repository, run python Web-LLM.py against a small query, and open the generated research session text file to see whether the saved content and links are good enough for your purpose.

## FAQ

### What exactly does an AI researcher do in Automated-AI-Web-Researcher-Ollama?

The README describes it as generating prioritized focus areas from your query, then for each area formulating search queries, running web searches, selecting pages, scraping them, and saving the content and source URLs into a research text file. It repeats with new focus areas derived from what it found until you stop it, then summarizes and answers questions about the findings.

### Who developed Ollama AI?

The README does not cover who develops Ollama. It only states that you install and configure Ollama following the instructions at ollama.ai, and that the researcher connects to it through the base_url in llm_config.py, which defaults to http://localhost:11434.

### Does Automated-AI-Web-Researcher-Ollama work on Windows?

The README says that to use it on Windows you should follow the instructions on the feature/windows-support branch, and that Linux and macOS use the main branch with the steps it lists.

### Which Ollama model should I use with Automated-AI-Web-Researcher-Ollama?

The README recommends picking a model with the required context length for lots of searches, naming phi3:3.8b-mini-128k-instruct or phi3:14b-medium-128k-instruct. You then set model_name in llm_config.py to the model you have set up in Ollama.

### How do I stop Automated-AI-Web-Researcher-Ollama during a research session?

The README lists q as the quit command, submitted by typing the letter and pressing CTRL+D. It also documents p to pause and assess progress, after which you enter c to continue or q to terminate, which produces a summary of the content collected so far.

## Sources

- [Issues](https://github.com/TheBlewish/Automated-AI-Web-Researcher-Ollama/issues)
- [License: MIT](https://github.com/TheBlewish/Automated-AI-Web-Researcher-Ollama/blob/main/LICENSE)
- [README](https://github.com/TheBlewish/Automated-AI-Web-Researcher-Ollama/blob/main/README.md)
- [TheBlewish/Automated-AI-Web-Researcher-Ollama on GitHub](https://github.com/TheBlewish/Automated-AI-Web-Researcher-Ollama)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/theblewish-automated-ai-web-researcher-ollama
