Model or dataset
guy-hartstein/company-research-agent avatar
guy-hartstein/company-research-agent

company-research-agent: a LangGraph pipeline that writes company reports with Gemini and GPT-5.1

An agentic company research tool powered by LangGraph and Tavily that conducts deep diligence on companies using a multi-agent framework. It leverages Google's Gemini 2.5 Flash and OpenAI's GPT-5.1 on the backend for inference.

2,294 stars322 forksPythonApache-2.0

At a glance

What is it?
guy-hartstein/company-research-agent splits company diligence across four research nodes, a curator, a briefing stage on Gemini 2.5 Flash and an editor stage on GPT-5.1. It needs four API keys before it will produce anything.
Who is it for?
Adopt it if you already pay for Tavily, Gemini and OpenAI and you want a self-hosted pipeline that turns a company name into a markdown report with a PDF export. Skip it if you want a single-key tool, if MongoDB persistence is a requirement rather than an option, or if you cannot accept that a failed run leaves no report behind.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 10 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap company-research-agent fills: four research passes, one report

Most company lookups end as a pile of open tabs. This project's answer is to split the work into named stages and let each stage do one thing. The README lists four research nodes, `CompanyAnalyzer`, `IndustryAnalyzer`, `FinancialAnalyst` and `NewsScanner`, each responsible for a different slice of the picture: core business, market position, financial metrics, and recent developments.

The target user is someone who needs a written brief on an unfamiliar company and is willing to pay per query across three providers to get it. That is not a consumer tool. It is closer to an internal research utility for a small team that already has Tavily, Gemini and OpenAI accounts, and that would rather run the pipeline on its own machine than paste a company name into a hosted product.

The trade-off is visible in the setup requirements. Four API keys are needed before the first report: Tavily, Google Gemini, OpenAI, and Google Maps. MongoDB is listed as optional. A tool that asks for four credentials is not competing on convenience. It is competing on the structure of the output.

How the LangGraph pipeline moves from search results to a formatted report

The architecture is a sequence of specialised nodes rather than one model with a long prompt. The four analyzers run first. `Collector` then aggregates what they returned. `Curator` applies content filtering and relevance scoring, and the README states that documents must clear a minimum Tavily relevance score, default 0.4, to proceed. Below that threshold a document is dropped, which means the quality ceiling of the whole report depends on what Tavily returns for the query in the first place.

After curation, `Briefing` generates category-specific summaries using Gemini 2.5 Flash. The README's stated reason is context handling: Gemini is used for high-context research synthesis and for maintaining context across multiple documents. `Editor` then compiles those briefings into the final report using GPT-5.1, handling deduplication, markdown formatting and real-time streaming. The README's justification for the split is that Gemini handles large context windows while GPT-5.1 follows exact formatting instructions more precisely.

That is a defensible design, and it is also the project's main cost driver. Every report pays for two model families. The backend is FastAPI with async support; research runs in the background and the React frontend polls `GET /research/{job_id}/report` until the report is ready. There is no websocket and no push notification, just polling against a job id.

Installing company-research-agent and running a first research job

The README recommends the setup script, which detects `uv` when present and falls back to standard Python tooling. Clone the repository and run it from the project root:

bash
git clone https://github.com/guy-hartstein/company-research-agent.git
cd company-research-agent
chmod +x setup.sh
./setup.sh

The script checks Python and Node.js versions, optionally creates a virtual environment, installs Python and Node dependencies, walks through environment variables, and can start both servers. If you prefer to control each step, the README gives the manual path: create a virtual environment with `uv venv .venv` or `python -m venv .venv`, activate it, then install dependencies with `uv pip install -r requirements.txt` or `pip install -r requirements.txt`, followed by `npm install` inside the `ui` directory.

Credentials go in a `.env` file at the project root. The README shows these keys for the backend:

env
TAVILY_API_KEY=your_tavily_key
GEMINI_API_KEY=your_gemini_key
OPENAI_API_KEY=your_openai_key

The README also lists a Google Maps API key and an optional `MONGODB_URI` for persistence, and `.env.example` ships the three keys above with empty values. The README documents the API surface as `POST /research` to submit a request, `GET /research/{job_id}/report` to poll for the finished report, and `POST /generate-pdf` to generate a PDF from report content. For a containerised run, `docker-compose.yml` starts the backend on port 8000 and a Node frontend on 5174 with `VITE_API_URL=http://localhost:8000`.

Where the pipeline breaks: the 0.4 threshold and the missing rollback story

The curator threshold is the sharpest limitation. A document scoring below 0.4 is discarded before any model sees it, and the README does not describe a fallback when a company simply has thin coverage. For a large public company with press coverage, that is fine. For a small private company, a regional subsidiary, or a newly incorporated entity, the analyzers may return almost nothing that clears the bar, and the editor will still produce a report. The failure mode is a confident-looking document built on a small set of surviving sources, not an error message.

There is a second gap. The README documents job submission and polling, but it does not document rollback, retry behaviour, or what happens to a job whose model call fails midway. Research runs asynchronously in the background, which is the right choice for a task that can take minutes, but it also means the failure surface is a job id that stops updating rather than an exception in your terminal.

A third constraint is cost opacity. The README does not publish token estimates, per-report pricing, or rate-limit handling for either provider. Since every run touches Tavily, Gemini and OpenAI, a batch of research jobs is a batch of metered calls across three bills. Anyone planning to run this at volume should instrument it themselves.

company-research-agent versus a general assistant with a search tool

The obvious alternative is a general-purpose assistant with web search, or a single-model agent built directly on LangChain. The difference is not model quality. It is the shape of the work.

A general assistant produces a conversational answer in one pass. company-research-agent produces an artefact: a stored report with a job id, a markdown structure, and a PDF export through `POST /generate-pdf`. The four-analyzer split also means the financial and news passes are separate calls with separate prompts, rather than whatever the assistant decides to search for. If your need is a quick verbal summary, the assistant is cheaper and faster. If your need is a repeatable document with the same sections every time, the pipeline structure is the point.

The second difference is hosting. A hosted assistant keeps your queries. This project runs locally or in your own container, and the README describes MongoDB persistence as optional, which suggests the default deployment keeps reports on the filesystem under `reports/`, mounted as a volume in the compose file. That matters if the companies you research are deal targets or clients.

Maintenance, licensing and what the Apache-2.0 header means here

The repository is not archived, and the last push was on 2026-09-08. The most recent release listed is v2.1.0 from 2026-07-04, following v2.0.0 in November 2025 and 2.0.1 in December 2025. The version history shows two major lines inside a year, which tells you the API surface has moved. The README documents `POST /research`, `GET /research/{job_id}/report` and `POST /generate-pdf`; those are the endpoints to check against your own copy before writing client code.

The licence is Apache-2.0, which permits commercial use and modification and includes an explicit patent grant. It also carries attribution and notice requirements, and the licence does not extend to the third-party services the pipeline calls. Your Tavily, Google and OpenAI usage stays governed by those vendors' terms, and the project's licence says nothing about them. That is a factual boundary, not legal advice.

Upgrade cost is mostly dependency churn. `requirements.txt` pins exact versions, including `langchain==1.3.9`, `langchain-openai==1.2.1`, `langchain-google-genai==3.0.3` and `tavily-python==0.7.24`. Pinned versions make installs reproducible but mean security or provider-API updates have to be applied by hand. The Dockerfile builds on `python:3.11-slim` and runs as a non-root `appuser`, so container rebuilds are the natural upgrade path.

Editorial conclusion

Adopt it if you already pay for Tavily, Gemini and OpenAI and you want a self-hosted pipeline that turns a company name into a markdown report with a PDF export. Skip it if you want a single-key tool, if MongoDB persistence is a requirement rather than an option, or if you cannot accept that a failed run leaves no report behind. Before committing, check that the curator threshold of 0.4 returns enough documents for the companies you care about, and confirm the exact model identifiers the briefing and editor nodes send to each provider.

Frequently asked questions

How do I create a research agent with company-research-agent?

Clone the repository, run `./setup.sh` or install `requirements.txt` manually, then fill in `TAVILY_API_KEY`, `OPENAI_API_KEY` and `GEMINI_API_KEY` in a root `.env` file. The README also lists a Google Maps API key and an optional MongoDB URI. The pipeline then runs the four analyzers, the collector, the curator, the briefing node and the editor.

Which models does company-research-agent use?

The README describes a dual model architecture: Gemini 2.5 Flash for high-context research synthesis in `briefing.py`, and GPT-5.1 for report formatting and editing in `editor.py`. The repository topics also reference gemini-3-flash, so the model identifiers are worth checking in your copy of the code.

Does company-research-agent need a Tavily API key?

Yes. `TAVILY_API_KEY` is one of the three keys in `.env.example`, and the curator scores documents using Tavily's relevance scoring with a default minimum threshold of 0.4. Without it, the search and curation stages have no source data.

Official sources

  1. guy-hartstein/company-research-agent on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/guy-hartstein-company-research-agent.svg)](https://hysenlabs.com/projects/guy-hartstein-company-research-agent)