PipesHub: an open-source, permission-aware context layer for workplace AI
PipesHub is an open-source platform for securely connecting enterprise knowledge to AI. Give AI agents trusted context and your team permission-aware search with verified citations across your business systems.
At a glance
- What is it?
- PipesHub connects Slack, Google Drive, GitHub, Microsoft 365, Notion and other business systems to AI, with block-level citations and source-level access controls. It is a self-hostable platform, not a library, and that shapes who should adopt it.
- Who is it for?
- Adopt PipesHub if you need a self-hosted context layer over several business systems and source-level permissions matter more than time to first answer. Skip it if you only need document Q&A over one folder, or if nobody on the team will operate a multi-service deployment.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem PipesHub solves is permissions, not retrieval
Most teams can already build a retrieval pipeline over a folder of documents. The hard part starts when the corpus is spread across Slack, Google Drive, GitHub, Microsoft 365 and Notion, and when the answer one employee gets must not contain a document another employee cannot open. That is the gap PipesHub targets. The README describes it as a platform for connecting AI applications to knowledge stored across business systems, and the two features it leads with are grounded answers with precise block citations and source-level access control enforcement. The audience is therefore an internal platform or IT team at a company large enough to have real permission boundaries, not an individual developer with a handful of PDFs. PipesHub also exposes that same context to external consumers: the README mentions APIs, SDKs, MCP tools and custom connectors, and the repository ships a Python SDK on PyPI, a Node SDK on npm and a separate Go SDK repository. So the product is both an end-user search and chat surface and a context backend that your own agents can call.
How the indexing and retrieval path is put together
The repository layout tells you more about the architecture than the feature list does. There is a backend/ directory in Python, a frontend/, a deployment/ directory, a separate integrations/ tree, plus integration-tests/ and loadtest/. The Dockerfile is split in two: Dockerfile.base holds what the comments call the slow layers (APT, Rust, Python dependencies, ML models and the runtime stack), while the top-level Dockerfile builds the app image from two registry images, pipeshubai/pipeshub-ai-base:python-deps and pipeshubai/pipeshub-ai-base:runtime. That split matters operationally. The base image bakes dependencies as of its publish time, and the app build reconciles with the current backend/python/pyproject.toml using uv, which skips already-satisfied packages. The Dockerfile comments also note that the published runtime base historically included LibreOffice Writer and Calc only, so PPT and PPTX previews could not be converted, and that CJK fallback fonts are installed until they appear in the published base. In other words, document conversion is done at request time inside the same image rather than by an external service. On the retrieval side the README claims graph-backed retrieval that captures relationships across enterprise data, plus multimodal handling of images, diagrams and scanned files. The mechanics of that graph are not spelled out in the README, so treat the claim as a direction rather than a documented design.
Installing PipesHub and running a first query
The README gives a single install path, a shell script served from get.pipeshub.com. It is the fastest way to see whether the platform fits, and it is the command the project itself puts in a callout box.
curl -fsSL https://get.pipeshub.com/install | bashThe repository also contains an install.sh at the top level, and a deployment/ directory, so the script is not the only route, though the README does not document what the script does step by step. If you prefer to build the image yourself, the Dockerfile documents build arguments for overriding the base images, including a slim variant that omits the pre-baked BGE model and is described as roughly 1.3 GB smaller, with the model downloading on first use.
docker build \
--build-arg PYTHON_DEPS_IMAGE=pipeshubai/pipeshub-ai-base:python-deps-slim \
-t pipeshubai/pipeshub-ai:slim .After the stack is up, the first real task is connecting one source. The README's connector section points at a connector gallery and the repository carries a CONNECTOR_INTEGRATION_PLAYBOOK.md for people writing their own. Connect a single system first (Slack or Google Drive are the examples the README uses), let indexing run, then ask a question in the chat surface and check that the answer carries block-level citations back to the original document. That citation check is the fastest way to confirm the pipeline is actually reading your content rather than answering from the model's own knowledge.
Where PipesHub stops being the right tool
The clearest limitation is operational weight. PipesHub is a multi-service platform with a Python backend, a separate frontend, a deployment directory and images that bake in machine learning models and a document conversion stack. The Dockerfile comments about LibreOffice components and CJK fonts are a good illustration: presentation previews depend on office tooling present in the runtime base image, and font coverage is a live concern that the project is patching around. If your team has no one to run and upgrade that, a smaller retrieval library over a single document store will serve you better. The second limitation is scope. PipesHub is built for enterprise systems with per-user permissions. If your corpus is a folder of public PDFs, permission-aware search buys you nothing and you pay for the connectors, the graph layer and the deployment surface anyway. Third, the README does not document rollback, backup or restore procedures, and it does not state which connectors enforce which permission model at the source. Those are exactly the details that decide whether a permission-aware deployment is trustworthy, and their absence from the README is a gap rather than a guarantee.
PipesHub compared with Glean, Onyx and Danswer
The comparison people actually search for is PipesHub against Glean, and the honest answer is that they are sold differently. Glean is a commercial workplace search product; PipesHub is Apache-2.0 and the README states it is fully self-hostable, with the claim that data never leaves your infrastructure and that you can bring any LLM provider. That is the real difference in approach: control of the deployment and the model endpoint, in exchange for running the infrastructure yourself. PipesHub Cloud is listed as coming soon with a waitlist, so today the open-source route is the route. Onyx and Danswer appear in the same search space as open-source enterprise search and assistants, and the practical distinction to check is connector coverage and how each project models source permissions, since that is where PipesHub puts its emphasis. PipesHub also ships an MCP package on npm, which matters if your plan is to feed existing agents rather than adopt another chat UI. None of these differences are settled by a feature table; they are settled by whether the specific connectors you need exist and behave.
Maintenance, upgrade cost and the Apache-2.0 licence
The last push to the default branch was on 2026-09-10, and the most recent release listed is v0.7.0 from 2026-08-26, following v0.6.0 on 2026-08-10 and a v0.6.0-beta on 2026-08-01. The cadence is roughly monthly on the 0.x line, which means minor versions can carry breaking changes and you should read CHANGELOG.md and the changelog/ directory before upgrading rather than assuming drop-in compatibility. Upgrading also means pulling new base images, since the app image is built from pipeshubai/pipeshub-ai-base tags; if you pinned a tag, the reconciliation step against pyproject.toml is what pulls in connector SDKs added since that base was published. Budget for that as part of every upgrade, not as an exception. The project is Apache-2.0, which permits commercial use and modification and includes an explicit patent grant, but it does not cover the third-party systems you connect to, your model provider's terms, or any hosted offering the project may launch. This is a description of the licence text, not legal advice; your counsel should review how the licence interacts with your distribution model.
Editorial conclusion
Adopt PipesHub if you need a self-hosted context layer over several business systems and source-level permissions matter more than time to first answer. Skip it if you only need document Q&A over one folder, or if nobody on the team will operate a multi-service deployment. Before committing, verify two things in your own environment: that the connectors you depend on actually honour the ACLs your sources enforce, and that your model endpoint works with the bring-your-own-model configuration the docs describe.
Frequently asked questions
Is there an open-source version of Glean available?
PipesHub is an Apache-2.0, self-hostable platform for connecting business systems to AI with permission-aware search and citations, which is the role Glean plays commercially. The README also lists PipesHub Cloud as coming soon, so the open-source deployment is the available route today.
What are the top alternatives to Glean?
PipesHub is one, alongside the other open-source enterprise search and assistant projects that appear in the same searches, such as Onyx and Danswer. The difference PipesHub emphasises is self-hosting with any LLM provider, so data stays in your infrastructure, plus MCP tools and SDKs for feeding your own agents.
Is PipesHub an open-source AI project?
Yes. The repository is public under the Apache-2.0 licence, the backend is written in Python, and the README describes it as an open-source platform for connecting AI applications to enterprise knowledge. It also publishes a Python SDK, a Node SDK and a Go SDK.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/pipeshub-ai-pipeshub-ai)