Inside NeuroAIHub/BrainPilot: a single-user runtime with a CPU-first sandbox build
BrainPilot: Automating Brain Discovery with Agentic Research
At a glance
- What is it?
- BrainPilot is an AGPL-3.0 agentic research workspace for brain science, built as fifteen npm workspaces with four compose files. The most informative parts of its own configuration are the ones that state limits: a runtime that is single-user by contract, a sandbox build that deliberately stops before the GPU stage, and case studies that publish their p values.
- Who is it for?
- Read this as a research prototype with a real engineering skeleton, not as a lab service. The configuration is unusually honest about its edges: one user per data root, a sandbox image whose GPU stage you have to ask for by name, a knowledge base builder compiled out at build time, and case studies that report p values of 0.083 rather than rounding them away.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The default sandbox build stops before the GPU stage on purpose
The default compose file pins a build target that exists to stop you paying for hardware you do not have. The sandbox image is built from docker/sandbox/Dockerfile, which is multi-stage, and its last stage is named gpu. A targetless build would therefore install the much larger GPU environment, so docker-compose.yml sets target: cpu and the comment above it says exactly that. GPU users are routed to a separate file, docker-compose.gpu.yml, rather than editing the default one. A third file, docker-compose.host.yml, overrides the network, and the default bridge network is described as isolated. A fourth variant, docker-compose.dynamic.yml, sits alongside them. Build args for HTTP_PROXY, HTTPS_PROXY, NPM_REGISTRY and APT_MIRROR hang off the sandbox service, so registry and package mirrors are a supported knob rather than a fork. For prebuilt images, setting BP_SANDBOX_IMAGE uses an official GHCR or ACR image without retagging, the distribution added in v0.1.2.
Two environment names resolve to one sandbox port
The sandbox container receives two names for the same port: PORT and AGENT_RUNTIME_PORT are both set from BP_SANDBOX_PORT, defaulting to 8081. The main service, backend plus web UI, is host-exposed on BP_MAIN_PORT, default 9001. The example environment file carries a comment warning that a conflict makes compose fail and telling you to change the value there, which is the only failure mode documented for the port pair. The same file spells out that the network is bridge by default and that host networking needs the override file named in the comment above it. One more variable, BP_MOCK, is passed straight through to the sandbox with no explanation of what it mocks, so its reach is not visible from the compose file alone. Everything else in that environment block is an Anthropic credential or a path, which makes the mock flag the one opaque entry in an otherwise explicit contract.
The runtime is single-user by contract, and the volume says so
The compose file is blunt about who the runtime is for: the runtime is single-user by contract, and multi-user hosting gives each user their own dataRoot or bind mount. Inside the container the data root is /root/.bp-root and the knowledge base root is a KnowledgeBase folder beneath it. The volume comment describes what that root holds, namely per-session directories named workspaces/<sid>/ and a cross-session persistent library at data/ that an issue reference says was flattened, so uploads and datasets placed there stay reusable across sessions instead of dying with the session that fetched them. That single paragraph is the entire persistence story, and it reads as a warning rather than a feature list. Anyone planning to serve several researchers from one host has to build that isolation outside this file, and the file offers no template for it.
A third-party gateway is a documented configuration, key forwarding included
The example environment file documents pointing the whole stack at a gateway you run. To target an Anthropic-compatible gateway you set both ANTHROPIC_BASE_URL and ANTHROPIC_MODEL, and the comment states that the key above is then sent as the gateway key rather than to Anthropic. Two numeric knobs are tunable: ANTHROPIC_CONTEXT_WINDOW, shown as 262144 against a default of 200000, and ANTHROPIC_MAX_TOKENS, defaulting to 8192. Anything more involved, meaning multiple providers, custom headers, openai-compatible endpoints or compatibility flags, means writing your own models.json, pointing BP_MODELS_JSON at it and naming BP_MODEL_PROVIDER, which defaults to the first provider in the file. models.example.json sits at the repository root as the starting shape. Read this section for one more reason: the v0.2.0 notes advertise model contexts configurable up to 1M tokens, while the documented override here stops at 262144.
The knowledge base builder is switched at build time, not at runtime
One flag decides whether the knowledge base builder appears, and it takes effect when the web bundle is built. VITE_KB_SETTINGS_ENABLED is set to 1 in the example file, with a comment explaining that a deployment provisioned with a knowledge base already can hide the Settings to Knowledge Base builder without disabling knowledge retrieval. Retrieval and the editing surface are therefore separate in practice: the builder is compiled in or compiled out, not toggled per user. The same pattern runs through the compose file, where most behaviour arrives as environment variables and only the build target, the network and the volumes are structural. Two concerns run the other way and stay runtime, the mock flag and the models file, both passed into the container as paths and values. For a managed deployment, that means the decision about who may edit the knowledge base is made at build time and cannot be revisited by editing an environment file.
Fifteen workspaces, seven of them in the typecheck script
The workspace list holds fifteen package paths, running from packages/protocol and packages/pi-sdk through packages/runtime, packages/backend-core, packages/web and four plugin packages, ending at packages/kb-scripts and packages/docs. The typecheck script names only seven: plugin-sdk, protocol, skills, runtime, backend-core, cli and client-cli. The web app, the auditor, got, research and monitor plugins, the knowledge base scripts and the docs site all sit outside that list. The build script compiles the Pi SDK first through a node script that also runs on postinstall, then builds every workspace with --if-present, which skips any package without a build script of its own. Testing is vitest at the root, with a dataset script that builds backend-core before exercising public dataset downloads. One script breaks the pattern: test:claude-mem runs a bash file, next to two other bash scripts that check upstream repositories.
Three releases on one day, and a tag dated after its announcement
The history is dense for a project that went open source on 2026-07-17 with v0.1.0. Within five weeks it shipped v0.1.1, v0.1.2, and then three versions on the single day of 2026-08-22: v0.2.0, v0.2.1 and v0.2.2. Two of those notes read as admissions. v0.2.1 makes Stop a hard boundary so cancelled model, tool, file and Trace work cannot resume after the user's next message. v0.2.2 aligns the hosted interface with currently available Cloud capabilities and removes unsupported control-plane requests, which says the interface had been asking for things the cloud side would not take. The version in package.json is 0.2.3, matching the newest tag, while the news entry announcing that release is dated 2026-09-08 and the release itself is stamped 2026-09-12. Around those releases sit the files a project of this size accumulates: SECURITY.md, RELEASING.md, CONTRIBUTING.md, a release-notes directory, a bilingual README pair and a Chinese locale of the same document.
The case studies publish the results that did not clear the bar
The case studies are the section worth reading, because they report the misses. On two-photon RSC calcium imaging with virtual-reality behavior, a five-part analysis workflow reached held-out Bayesian decoding at MAE = 16.8 cm and r = 0.646. Across 58 Allen Neuropixels sessions, three functional measures correlated positively with anatomical hierarchy and none crossed the conventional significance threshold, at p = 0.083, 0.243 and 0.058. A frozen 279-region pain-connectivity signature assigned a higher pain response in 9 of 10 held-out subjects and transferred unevenly across the Japan and UK cohorts, at AUC = 0.793 and 0.699. The fourth case names BCI Competition IV 2a for EEG motor imagery, and its text ends there, part way through a word. The process side is documented too, with the Graph of Trace published as an ACL 2026 demo and an evaluation file committed at the repository root under a date of 2026-08-05.
Editorial conclusion
Read this as a research prototype with a real engineering skeleton, not as a lab service. The configuration is unusually honest about its edges: one user per data root, a sandbox image whose GPU stage you have to ask for by name, a knowledge base builder compiled out at build time, and case studies that report p values of 0.083 rather than rounding them away. Before you plan anything around it, confirm four things. Whether single-user by contract is acceptable for your setting, because multi-user hosting means a data root and bind mount per person and nothing in the compose file does that for you. Whether the license fits your deployment, since the repository is AGPL-3.0-only while the project also points at a hosted site and Cloud capabilities. Which model context you actually need, given the environment example documents 262144 against a 200000 default while the release notes advertise up to 1M. And whether the gateway configuration sends your key somewhere you intend, because the example file states plainly that the key is forwarded as the gateway key once ANTHROPIC_BASE_URL is set.
Frequently asked questions
What is NeuroAIHub/BrainPilot?
An open-source agentic research workspace for brain science, released under AGPL-3.0. A Principal Investigator agent talks to the user and coordinates specialist agents including a librarian, experimentalist, engineer, writer and auditor, and each session is recorded as a Graph of Trace so intermediate actions, evidence and claims can be inspected.
How do I run BrainPilot with Docker?
The default docker-compose.yml builds the sandbox with an explicit cpu target and exposes the main service on BP_MAIN_PORT, 9001 by default, with the sandbox on BP_SANDBOX_PORT, 8081 by default. GPU builds come from docker-compose.gpu.yml, host networking from docker-compose.host.yml, and setting BP_SANDBOX_IMAGE lets you use an official GHCR or ACR image without retagging it.
Can BrainPilot talk to a third-party model gateway?
Yes, by setting both ANTHROPIC_BASE_URL and ANTHROPIC_MODEL, and the example environment file states that ANTHROPIC_API_KEY is then sent as the gateway key. For multiple providers, custom headers or openai-compatible endpoints you write your own models.json, point BP_MODELS_JSON at it, and optionally set BP_MODEL_PROVIDER, which defaults to the file's first provider.
Is BrainPilot a multi-user application?
The compose file states that the runtime is single-user by contract, with per-session workspaces under the data root and a cross-session library alongside them. Multi-user hosting gives each user their own dataRoot or bind mount, which the deployment files do not set up for you.
What changed in BrainPilot v0.2.1 and v0.2.2?
v0.2.1 makes Stop a hard boundary, so cancelled model, tool, file and Trace work cannot resume after the user's next message. v0.2.2, released the same day, aligns the hosted interface with currently available Cloud capabilities and removes unsupported control-plane requests.
Which datasets do the BrainPilot case studies use?
Two-photon RSC calcium imaging with virtual-reality behavior, where held-out Bayesian decoding reached MAE = 16.8 cm and r = 0.646, plus 58 Allen Neuropixels sessions, a frozen 279-region fMRI pain-connectivity signature that transferred at AUC = 0.793 and 0.699 across the Japan and UK cohorts, and BCI Competition IV 2a for EEG motor imagery.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/neuroaihub-brainpilot)