Open Glean: an AI workspace you can read end to end
An open-source AI platform for knowledge work. Connect your apps, find answers, and get work done.
At a glance
- What is it?
- Hydra DB publishes the whole server for its Open Glean assistant: a Next.js 16 app that streams cited answers over a document store, keeps API keys in encrypted cookies, and hands you the research pipeline in the open. Apache 2.0, version 0.1.0, no tagged releases yet.
- Who is it for?
- Open Glean is worth reading if you are building a retrieval assistant on a document store, because the parts that are normally hidden, the DAG planner, the session encryption, the DocumentDB workaround, the concurrency cap, are all in the repository rather than behind a product wall. It is not ready to adopt blindly: version 0.1.0, private in package.json, no tagged releases, two open issues, and one documented dead end in the IAM authentication path.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 19 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Hydra DB built and why it gave away the server
The repository description is short: an open-source AI platform for knowledge work, connect your apps, find answers, and get work done. The homepage is hydradb.com, the license is Apache-2.0, and the language is TypeScript. At 1531 stars and 518 forks with two open issues, the interest is well ahead of the maintenance queue, which tells you something about where this sits in the company's plan.
The framing in the README is more specific. Open Glean is the AI workspace over Hydra DB, and the promise is that you ask a question across your memories, files and connected apps, and the system retrieves the context, writes the answer, and cites its sources. The grant is Apache-2.0, and there is a NOTICE file that carves out connector logos as third-party trademarks. That is a careful carve-out, and it says something about which half of the product is commercially sensitive.
The interesting decision is the shape of the giveaway. This is not a client library or a thin wrapper around an API. It is the whole Next.js server, including the proxy routes, the encryption of session cookies, the Dockerfile, and an AGENTS.md that documents conventions for agents working in the tree. The README is explicit about the boundary: the Next.js server proxies requests to the official `@hydradb/sdk`, and keys are held in an encrypted, httpOnly session cookie used server side, never in the browser. You can also run it yourself with your own Hydra DB key, which means the hosted version is a convenience rather than the only path.
One detail to check before anyone reads too much into the version number. package.json declares version 0.1.0 and private true. There are no tagged releases in the repository at all, and the last push to the main branch was 2026-09-17. What you have here is a source snapshot of an actively developed application, not a published package with a changelog and an upgrade path.
The features list is short, and one item is the whole product
The README lists seven features. Ask is one composer that retrieves context from your Hydra database and streams an answer with inline citations, a sources panel, and optional web search through OpenRouter's web plugin. Context collects memories, files, saved pages and connector-synced knowledge in one place. Collections scope a query to one collection, treated as a sub-tenant. Scope switching picks databases and collections from the top bar, and retrieval fans out across every selected one. Integrations connects Hydra's connectors, verifies credentials, discovers resources and starts syncing without leaving the app. Mindmap renders the knowledge graph Hydra builds from your context. Bring your own model accepts any OpenAI-compatible endpoint, with a searchable OpenRouter picker and favourites.
Of those, Deep Research is the one that would justify its own section in your notes, because the mechanism is described with unusual precision. For questions a single query cannot answer, Open Glean plans a directed acyclic graph of sub-questions, runs each level against Hydra in parallel, writes a finding per branch, dedupes sources into one numbered citation list, then writes the answer. A timeline shows the plan and the live progress. That is a concrete description of a research pipeline rather than a marketing sentence, and it is the part most worth borrowing if you are building the same thing.
The graph rendering is also where the dependency list stops looking ordinary. Alongside the expected Next.js, React and Tailwind entries, package.json carries d3-force, react-force-graph-2d and framer-motion. A force-directed knowledge graph on a canvas inside a server-rendered app is exactly where those two libraries stop being decoration and start being the feature. The rest of the dependency list is short for a project of this size: `@hydradb/sdk` at 2.1.2, the mongodb driver, `@aws-sdk/rds-signer`, clsx and Phosphor icons. That is a build graph dominated by the framework rather than by a large third-party surface, which is a good sign for how much of the interesting logic is the project's own.
Node and TypeScript are pinned tightly, with next at 16.3.3, react and react-dom at 19.2.8 and TypeScript 5 on the dev side. Exact pins rather than ranges, in a repository with no releases, mean you inherit a specific combination and you own the upgrade.
How keys are handled, and why the pinning rule exists
The security section of the README is the most specific documentation in the repository, and one rule in it deserves understanding rather than copying.
API keys are not stored in the browser. When you connect in Settings, the key is verified server-side and stored in an AES-256-GCM-encrypted, httpOnly session cookie, and every proxied request resolves it on the server. Nothing key-shaped is inlined into the JavaScript bundle, into localStorage, or into any client state. The encryption is keyed by `OPEN_GLEAN_SESSION_SECRET`, which the README marks as required in production and describes as any random string of 16 or more characters. The Dockerfile handles the build-time wrinkle by setting a placeholder value during the build, with a comment explaining that the real secret is supplied when the container runs.
The rule that deserves attention is endpoint pinning. A stored key is pinned to the endpoint it was stored with, and a request may only choose the LLM base URL when it supplies its own API key. Otherwise, as the README puts it, a caller could pair an attacker-controlled URL with the server's key and have it sent there in an Authorization header. That is a confused deputy problem, and it is easy to introduce by accident in any app that accepts both a stored credential and a caller-supplied base URL. A related flag, `OPEN_GLEAN_ALLOW_PRIVATE_LLM_URL`, exists so a base URL can sit on a private or loopback address for something like Ollama or LM Studio. It is off by default, and even when enabled it permits plaintext http only on a private address, still requires https for a public host, and rejects non-http schemes entirely.
For a chat product, that combination is well judged. Server-side verification, encrypted cookies, no client persistence, and a scheme check on the one endpoint that has to stay flexible is more care than most projects of this size take. The session secret being the only required variable also means the minimum viable configuration is one string, which is a good property for something people will try before they trust it.
Persistence has three shapes, one of which is a workaround
Conversations persist to MongoDB when it is reachable, with a localStorage fallback, and non-secret settings stay in the browser. The example environment file spells out the options, and it is worth reading because the deployment story is broader than a single connection string.
Option A is direct MongoDB, either a local instance for development or a hosted cluster, with the database name defaulting to open_glean. Option B is a proxy path described as Vercel to API Gateway to Lambda to DocumentDB, where the URI becomes a bare HTTPS execute-api endpoint and `MONGODB_PROXY_KEY` carries the shared secret. The Lambda reads the same value under the name PROXY_KEY, and the README states plainly that the two names must match. That is the kind of detail that costs an afternoon when it is wrong and saves a day when it is documented.
The DocumentDB story comes with an explicit warning that is unusual to find in a README. DocumentDB does not support IAM database authentication. The IAM token path in the code cannot authenticate against Amazon DocumentDB, and the AWS variables that path reads are documented only to describe what it expects. For DocumentDB, the recommendation is a standard MongoDB connection string or the Lambda proxy. Since the RDS signer is in the dependency list, the code path exists and is reachable, and the documentation is telling you not to point it at that particular managed service. A README that says this about its own feature is worth more than a working demo would have been.
There is also a deep-research concurrency cap, `OPEN_GLEAN_MAX_CONCURRENT_RESEARCH`, defaulting to 3 per instance, with the README noting that each run costs many LLM and retrieval calls. It is the kind of variable that tells you where the real cost sits. In a young application that has never shipped a release, an explicit cost ceiling is a reasonable sign that the operational thinking ran ahead of the release discipline.
The build, the container, and the parts you inherit
Local startup is two commands after you set one variable. The example file is a genuine catalogue rather than a stub: it opens with the required session secret, then Hydra settings including the optional deployment-level shared key, the base URL, and an optional default database so that shared-key deployments do not silently miss data. Then come the LLM provider settings, the chat persistence block, and a note that the model key is optional and only affects answer synthesis.
npm ci
npm run devThe Dockerfile is the clearest statement of what a deployment looks like. It is a three-stage build on node:24-slim: install and build with the full toolchain, then copy only the standalone server into a small runtime, which is what the next config's standalone output mode produces. The runtime stage copies the public directory, the standalone output and the static assets, plus a certs directory whose comment explains that the DocumentDB CA bundle is read at runtime by the Mongo library when a cluster endpoint is configured, and that it is a public certificate safe to ship. It creates a system user with uid 1001, switches to it, exposes port 3000 and runs the standalone server.
The scripts in package.json are conventional and complete: dev, build, start, lint, typecheck, test and test:watch, with tests on vitest and a config file named for it. TypeScript 5, ESLint 9 with the Next config, and Tailwind v4 through the PostCSS plugin, alongside an eslint config and a tsconfig at the top level. The type and lint gates are real rather than aspirational, and a test runner is configured even though no test files are visible in the tree listing.
The rest of the tree is where the shape of the application shows: an app directory, a components directory, a lib directory, a lambda directory, a certs directory, and a proxy file at the root. There are also AGENTS.md, CLAUDE.md, CODEOWNERS, CODE_OF_CONDUCT.md, CONTRIBUTING.md and SECURITY.md. Two agent instruction files in one repository is a deliberate choice, and it is consistent with the framing of the project as an assistant for knowledge work. It also means the conventions are written down, which is the single most useful thing for anyone contributing.
Where the gaps are, and what to check first
Three things deserve attention before anyone builds on this. The first is maturity signalling: version 0.1.0, private in package.json, and no tagged releases, so there is no upgrade path and no changelog to diff against. The pinned framework versions mean you inherit a specific combination rather than a range, and you own the upgrade when Next moves.
The second is the name. Search results for this project collide with a commercial enterprise search product that shares the word Glean. If you are evaluating it, read the README rather than the search results, because the pricing and feature comparisons that dominate those results describe a different company with a different architecture. That is not a criticism of the project. It is a practical warning about the search surface, and it is why the self-hosting path matters here more than it would for a better-named project.
The third is the model layer. The README states there is no built-in default for the model variable, and that without it or a per-user model, answers are unavailable. The example environment file names two current candidates with their context windows, a Gemini flash variant at 1M context and a DeepSeek flash variant at 1.3M context, and repeats that a private or loopback base URL can be enabled for a local model. Large context windows matter here for a reason beyond chat. Deep Research fans retrieval out across every selected collection, then has to hold sub-question findings and one deduped citation list at the same time, and web search citations have to be merged with the Hydra sources in a single answer without losing the numbering.
So the evaluation order is straightforward. Run it against your own Hydra DB key, add one model, ask something whose answer you can verify, then ask something that forces Deep Research and watch the timeline. After that, read the Mongo library for the DocumentDB endpoint detection and the lambda directory for the proxy. If the answer quality holds up and the pipeline is as legible as the documentation claims, you have a reference implementation and a set of patterns worth lifting.
Editorial conclusion
Open Glean is worth reading if you are building a retrieval assistant on a document store, because the parts that are normally hidden, the DAG planner, the session encryption, the DocumentDB workaround, the concurrency cap, are all in the repository rather than behind a product wall. It is not ready to adopt blindly: version 0.1.0, private in package.json, no tagged releases, two open issues, and one documented dead end in the IAM authentication path. Run it against your own Hydra DB key before committing. If your need is a hosted enterprise search product with connectors and permissions already handled, buy that instead. If your need is to see how a cited-answer assistant is actually assembled, this is a rare chance to read one.
Frequently asked questions
What is Open Glean and who maintains it?
Open Glean is the open-source AI workspace that Hydra DB publishes for its hosted assistant. The repository is hydra-db/open-glean, written in TypeScript, licensed Apache-2.0, with the homepage at hydradb.com. It streams cited answers across your memories, files and connected apps, and it can be self-hosted with your own Hydra DB key rather than only used through the hosted service.
Is Open Glean free to run yourself?
The code is Apache-2.0, so running it yourself costs only infrastructure. In practice you need three things: a Hydra DB API key, a session secret of 16 or more characters in OPEN_GLEAN_SESSION_SECRET, and optionally an LLM key. Direct MongoDB or the Lambda proxy handles chat persistence, and web search runs through your LLM provider's web plugin.
How does Open Glean handle Deep Research?
For a question a single query cannot answer, it plans a directed acyclic graph of sub-questions, runs each level against Hydra in parallel, writes one finding per branch, dedupes the sources into a single numbered citation list, then writes the answer. A timeline shows the plan and live progress. Concurrency is capped at 3 runs per instance by default through OPEN_GLEAN_MAX_CONCURRENT_RESEARCH.
Where are API keys stored in Open Glean?
Keys are verified server-side when you connect in Settings, then stored in an AES-256-GCM-encrypted httpOnly session cookie keyed by OPEN_GLEAN_SESSION_SECRET. Nothing key-shaped reaches the JavaScript bundle, localStorage or any client state. A stored key is also pinned to the endpoint it was stored with, so a caller cannot pair the server's key with an attacker-controlled LLM base URL.
Can Open Glean use Amazon DocumentDB for chat storage?
Yes, but not through IAM database authentication, which the README states DocumentDB does not support. The IAM token path in the code cannot authenticate against DocumentDB, so for that service use a standard MongoDB connection string or the Lambda proxy path, where the URI points at an execute-api endpoint and MONGODB_PROXY_KEY matches the Lambda's PROXY_KEY.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hydra-db-open-glean)