# BambooAI replays its own report, and two dependency files disagree

> A self-hosted LLM analyst that runs Python in a persistent kernel and re-executes the cells a report cites in a fresh kernel before showing it, packaged at 2.0.2 while the newest tag is v2.0.0, with two dependency manifests that list different things.

**pgalko/BambooAI** — A Python library powered by Language Models (LLMs) for conversational data discovery and analysis.

- Repository: https://github.com/pgalko/BambooAI
- Stars: 792 · Forks: 87
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/pgalko-bambooai

## The report is replayed in a fresh kernel before it is shown

The verification step is mechanical. Once the analyst issues REPORT the run ends, and the cells the report cites are assembled in execution order into a script, which is then run in a fresh kernel against the original data. The numbers the report quotes are compared with that fresh output, and a closing card states the result in the form of a count, for example replay reproduced 41 numbers in a fresh run. The figures shown in the Plots tab are captured from the reproduction rather than from the first pass, so a chart you see is the one the replay drew. The limit is structural: only cited cells are replayed, so a number the analyst left out of the report is not checked by anything. The one page contract at analyst/contract.md, which is the same for every model, asks for a best estimate with its uncertainty, the conditions it holds under, how it was established, and the cells and figures it rests on, usually one to three Plotly figures.

## The turn budget is a preset, and only one preset's numbers are published

Each mode sets both a turn budget and a dollar budget that the analyst can see while it works, and both come from a preset named tier_properties in the configuration. The page gives the values for one preset only, the one called performance: Quick gets 4 turns, Deep gets 15, and Adaptive gets 50 with a self-review every 8 turns. Adaptive is the one that can change course, because the review turns run on the Reviewer seat, described as a stronger model, which reads the run so far and is able to redirect it. A 50 turn Adaptive run therefore admits roughly six redirections from a different model than the one doing the work. The names of the other presets and their budgets are not on the page, so a reader cannot tell what the cheaper settings actually cost in turns.

## Ten providers are named and seven client libraries are installed

The feature line claims ten providers with per-seat configuration: OpenRouter, OpenAI, Anthropic, Google, Groq, Mistral, xAI, DeepSeek, Ollama and vLLM. The dependency list tells a narrower story. It installs client packages for OpenAI, Anthropic, two separate Google packages, Groq, Mistral and Ollama, and nothing for xAI or DeepSeek, both of which are reachable over an OpenAI compatible endpoint instead. vLLM is a server rather than a client library, so it belongs in the same group. Two entries carry version floors, openai at 1.30 or newer and ollama at 0.5 or newer, while the rest float. Seats are what make the provider count matter: analyst, reviewer, rewriter and knowledge distiller each carry their own model and reasoning effort, so a single Adaptive run can be billed across several providers at once.

## Two dependency manifests list different things

The packaging metadata is explicit about scope. A comment above the dependency block says these are the self-hosted edition's dependencies, points at docs/OSS_DESIGN.md, and adds that nothing of the hosted service is here. The separate requirements.txt in the repository then lists supabase, stripe and google-cloud-storage, which are hosted service concerns, alongside selenium, python-jose with cryptography, and kaleido. The same file also installs geopandas, yfinance and kaleido unconditionally, while pyproject puts geopandas, yfinance and kaleido into an optional extra named analysis that you install only when the work needs them. Neither file pins versions. So a deployment that installs requirements.txt gets a larger dependency set than the packaging metadata describes, including clients for infrastructure the metadata says is absent.

## The package version runs ahead of the newest tag, and a different readme ships

The version field in pyproject.toml reads 2.0.2. The newest GitHub release is v2.0.0, tagged on 2026-09-16, and the two tags before it are v0.4.26 from 2025-10-31 and v0.4.25 from 2025-08-10, so the visible history jumps from the 0.4 line straight to 2.0 with no 1.x tag in between. The consequence for anyone pinning: the artifact on PyPI can be two patch releases ahead of anything tagged, and nothing on the page maps 2.0.1 or 2.0.2 back to a change. The readme has its own split. The packaging points at docs/PYPI_README.md rather than README.md, with a comment explaining that PyPI cannot resolve the repository-relative images of the main file, so the description you read on the package index is a different document from the one on the repository.

## Docker is a hard requirement, and the code runs in a container inside it

The requirements are Python 3.11 or newer and a running Docker, Docker Desktop on macOS and Windows and Docker Engine on Linux, and the page says outright that Docker is where the model-written code runs. The isolation is the point rather than a side effect: the executor container is built from the Dockerfile inside the package and is reachable only from your machine, and the summary line says the code the model writes runs in a container, not under your account. Three install routes are offered. The recommended one is two commands:

```bash
pip install bambooai
bambooai serve
```

The second builds from the repository instead, with a clone, a virtual environment and an editable install. The third is the hosted service at bambooai.org, described as the same analyst with nothing to install. On first use, serve creates a working folder at ~/bambooai, so a headless server with no Docker gets the package but not the analyst.

## Kernel state survives a failed cell, not a new question

There are two separate kinds of persistence. Inside a run, the analysis is a sequence of cells in a persistent kernel with a checkpoint per cell, and a failed cell is rolled back so the namespace is exactly as it was before the attempt. Across a conversation, a follow-up question continues with the variables the previous run left behind, while a new question starts fresh, and runs form chains you can step through. Across sessions, saving a run hands the transcript to the Knowledge Distiller seat, which writes a card about the dataset into memory/<user>/memory_pack.yaml, covering what the columns mean in practice, what to watch for and what worked. The RECALL action reads that card back in later runs. Note the file is keyed by user inside a local folder, so what the analyst carries forward is a local artefact rather than anything on the account.

## Seven actions, and one of them blocks the run on a person

Each turn takes exactly one action from a fixed set of seven, and the printed output of a cell becomes the next turn's input. CELL runs one Python cell. SHOW n re-opens the full output of an earlier cell, which matters because turns read truncated outputs. NAMES lists the variables the kernel holds. RECALL consults the memory pack. SEARCH performs a web search with Google AI grounding, summarised and cited, and needs a Gemini key. ASK puts one question to the person and waits for the answer. REPORT writes the report and ends the run, which is what triggers the replay. ASK is the one that changes the shape of a run, because the run cannot continue until a human replies, and the page does not say whether waiting counts against the turn budget or whether there is any timeout on it.

## Conclusion

The replay step is the part worth trusting and the packaging is the part worth checking. Replayed numbers are verified against fresh output, but only for the cells the analyst chose to cite, so a report can still omit a number it got wrong. Before installing, confirm Docker is available, because the model-written code runs in a container built from the package, and read pyproject.toml rather than requirements.txt, since the second file lists hosted-service dependencies the packaging metadata says are not part of this edition.

## FAQ

### Does BambooAI need Docker to run?

Yes. The page lists Docker as a requirement, Docker Desktop on macOS and Windows or Docker Engine on Linux, because the code the model writes executes inside an executor container built from the Dockerfile shipped in the package. Python 3.11 or newer is the other stated requirement.

### What does BambooAI do before showing a report?

It assembles the cells the report cites into a script in execution order, runs it in a fresh kernel against the original data, and compares the quoted numbers with the fresh output. A closing card reports how many reproduced, and the figures in the Plots tab come from that reproduction run.

### Which model providers does BambooAI support?

Ten are named: OpenRouter, OpenAI, Anthropic, Google, Groq, Mistral, xAI, DeepSeek, Ollama and vLLM. Each seat, including the analyst, reviewer, rewriter and knowledge distiller, carries its own model and reasoning effort, and local models keep everything on the machine.

### How long can a BambooAI run take?

Each mode carries a turn budget and a dollar budget the analyst can see, taken from a preset in the configuration. At the preset named performance, Quick gets 4 turns, Deep gets 15, and Adaptive gets 50 with a self-review every 8 turns on a stronger reviewer model.

### Is the BambooAI pip package the same as the hosted service?

The packaging metadata says the dependency block is the self-hosted edition's and that nothing of the hosted service is in it, while the hosted version at bambooai.org runs the same analyst as a service. A separate requirements.txt in the repository does list supabase and stripe.

## Sources

- [Issues](https://github.com/pgalko/BambooAI/issues)
- [License: MIT](https://github.com/pgalko/BambooAI/blob/main/LICENSE)
- [pgalko/BambooAI on GitHub](https://github.com/pgalko/BambooAI)
- [README](https://github.com/pgalko/BambooAI/blob/main/README.md)
- [Releases](https://github.com/pgalko/BambooAI/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/pgalko-bambooai
