xynehq/xyne: a self-hosted, permissions-aware search and answer engine for Google Workspace
AI-first Search & Answer Engine for work. Open-source alternative to Glean.
At a glance
- What is it?
- Xyne indexes Google Workspace and answers questions with sources, runs from a docker-compose file, and takes any LLM. The repository README now points readers to a new home, which changes what adopting it means.
- Who is it for?
- Adopt Xyne if you want a self-hosted answer engine over Google Workspace data and are willing to run Vespa, Bun and an LLM endpoint yourself. Do not adopt this repository if you need a project with an active upstream here; the README states it has moved to github.com/juspay/xyne-spaces, and the last push was on 2026-09-10.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Xyne solves, and who it is actually for
Work information sits in separate systems: documents, files, repositories, projects, issues, email, chat messages and tickets. Xyne's README frames the problem as fragmentation across SaaS apps and positions the product as an AI-first search and answer engine that connects to those apps, indexes the data and maps a graph of relationships, so a user gets a Google plus ChatGPT style experience with sources attached to each answer. The stated audience is an organization, not an individual: permissions-aware access is listed as a feature, meaning the engine enforces the permissions that already exist in the connected apps rather than granting a shared pool of content. The README also names the integrations in scope. Google Workspace is described as live across Drive, PDFs, docs, sheets, slides, calendar, contacts and Gmail including attachments, and the README claims most enterprise search products mean only Drive when they say Google Workspace. Slack and Atlassian are listed as next, not as shipped. If your content lives mainly in Confluence, Jira or Slack, this repository is not yet the tool for that job. If your organization runs on Google Workspace and wants answers with citations without sending documents to a vendor, the shape of the product matches the need.
How the pieces fit: Bun, Vespa and a graph of relationships
The top-level layout separates a server directory and a frontend directory, with a Dockerfile that builds both. The Dockerfile starts from oven/bun:1, copies server/package.json and server/bun.lock alongside frontend/package.json and frontend/bun.lockb, runs bun install in each, then builds the frontend with bun run build. That tells you the runtime is Bun rather than Node, and that the frontend is compiled at image build time. The same Dockerfile installs canvas-related system libraries (libcairo2-dev, libpango1.0-dev, libjpeg-dev, libgif-dev, librsvg2-dev, libpixman-1-dev, libfontconfig1-dev, libfreetype6-dev) and downloads the Vespa CLI tarball at version 8.453.24 into /usr/local/bin. So Vespa is part of the storage and retrieval path, and the image pins a specific CLI build rather than tracking latest. A separate Dockerfile-vespa exists at the top level, which suggests the search backend can be built or run independently of the application image. The README describes the data flow in one sentence: connect to applications, index the data securely, map a graph of relationships, then answer with up-to-date sources. The ingestion side is described as multi threaded, which matters because the first sync of a Workspace tenant is the expensive step. Nothing in the README documents the schema of the relationship graph or how permissions are evaluated at query time, so that part has to be read from the server source or the documentation site.
Installing Xyne: docker-compose, then a Service Account
The README states Xyne can be run locally or deployed on a virtual machine with a single docker-compose command, and points to a Quickstart page in the documentation for local use and an AWS EC2 page for cloud deployment. The repository itself carries start.sh and start-aws-linux.sh at the top level, plus a deployment directory, which is where the compose files and startup scripts live. The README does not reproduce the compose file, so the exact service names, ports and environment variables have to come from the docs or from deployment/. What the README does give is the Google Workspace setup path: follow the Service Account integration steps in the documentation. The Dockerfile shows that the container expects production mode, set through the environment:
ENV NODE_ENV=productionBefore that, the image installs the Vespa CLI at a pinned version, which is the piece most likely to surprise a first-time deployer:
wget https://github.com/vespa-engine/vespa/releases/download/v8.453.24/vespa-cli_8.453.24_linux_amd64.tar.gz
tar -xzf vespa-cli_8.453.24_linux_amd64.tar.gz
mv vespa-cli_8.453.24_linux_amd64/bin/vespa /usr/local/bin/A first real use is then a workspace question rather than a keyword query: connect the Service Account, wait for the initial index, and ask something that spans systems, for example a customer name that appears in Gmail threads and in Drive documents. The README's own examples are finding a file, triaging an issue, asking about a customer, deal, feature or ticket, and discovering the right people. If the answer returns without a source link, the integration is not the problem; the index has not caught up.
The repository has moved, and that is the first thing to check
The README opens with a notice that the repository is no longer maintained here and names github.com/juspay/xyne-spaces as the new home. That single line reframes every other fact on this page. The last push to this repository was on 2026-09-10, and the most recent release listed is v3.1.1 from 2025-09-30, so the code here is not abandoned in the sense of being frozen years ago, but the project's own documentation tells you the upstream has shifted. For an adopter, the practical consequence is that issues, pull requests and future releases may land in a different repository than the one you cloned. A pinned commit here will keep working, but fixes for Vespa compatibility or Google API changes may only appear at the new location. Anyone evaluating Xyne should read both repositories before deciding, and should treat this README as a pointer rather than a statement of current development. The Apache-2.0 licence travels with the code, so the fork itself is not the risk; the risk is building an internal deployment on a tree that no longer receives the project's attention.
Where Xyne is the wrong tool
The README is explicit that the shipped integration is Google Workspace, with Slack and Atlassian marked as next. An organization whose knowledge lives in Confluence, Jira and Slack would be adopting a search engine with nothing to search. The second boundary is operational. Xyne is self-hosted by design, which the README lists as a feature, and that means you run Vespa, Bun, the server and the frontend, plus an LLM endpoint. The README says the model layer is agnostic and can point at any LLM, including a local DeepSeek via ollama, and that no telemetry is collected and no training happens on your data or prompts. Those are real properties, but they shift the burden of capacity planning, upgrades and backup onto your team. A hosted product such as Glean or Gemini, which the README names as the alternatives it is positioned against, removes that burden and charges for it. The third boundary is permissions. The README states that access control is enforced live against the apps' existing permissions, which is the correct design for enterprise search, but it also means a misconfigured Service Account can silently narrow or widen what the index contains. There is no documented dry-run mode for checking what a given account can see before the first full sync.
Alternatives and the difference in approach
The README positions Xyne against Glean, Gemini and MS Copilot. The difference is not the feature list but where the index and the model call happen. Glean is a hosted enterprise search product: the index lives with the vendor and the connectors are maintained by the vendor. Gemini and Copilot are tied to a specific cloud and, in Copilot's case, to the Microsoft 365 graph. Xyne takes the opposite route. The index runs on your infrastructure, the LLM is a configurable endpoint, and the connectors are code in a repository you can read. For an organization with a data residency constraint or an internal model endpoint, that is the deciding factor. For an organization without an operations team, it is the disqualifying one. A second, less obvious alternative is doing nothing and relying on the search box inside each SaaS app. That works until a question spans two systems, which is exactly the case the README describes when it talks about asking everything about a customer, deal, feature or ticket. The honest comparison is therefore not Xyne versus Glean on features; it is self-hosted retrieval with your own model endpoint versus a vendor-hosted index, and the cost sits in different columns either way.
Maintenance, licence and what an upgrade costs you
Xyne is Apache-2.0. That permits commercial use, modification and redistribution, and it includes a patent grant, but it also means the project offers no warranty, and this article is not legal advice; if you redistribute a modified build, read the licence text and your own obligations. On maintenance, the facts are narrow and worth stating plainly: the README says the repository has moved, the last push here was on 2026-09-10, and the newest release listed is v3.1.1 from 2025-09-30. Upgrade cost is dominated by two moving parts visible in the repository. The first is Vespa: the Dockerfile pins the CLI at 8.453.24, so a server running a different Vespa version is an untested combination, and Vespa upgrades generally require planning rather than a container restart. The second is Bun: the image is built from oven/bun:1, and the frontend lockfile is frontend/bun.lockb while the server uses server/bun.lock, so both dependency trees move independently. There is also a renovate.json at the top level, which indicates automated dependency updates are configured, but the README does not document a rollback procedure or a supported upgrade path between releases. Anyone running this in production should treat the deployment directory as the source of truth for how a version is started, and should keep the Vespa version and the CLI version in step.
Editorial conclusion
Adopt Xyne if you want a self-hosted answer engine over Google Workspace data and are willing to run Vespa, Bun and an LLM endpoint yourself. Do not adopt this repository if you need a project with an active upstream here; the README states it has moved to github.com/juspay/xyne-spaces, and the last push was on 2026-09-10. Before deploying, verify the current home, confirm the Vespa CLI version pinned in the Dockerfile (8.453.24) matches the server you intend to run, and confirm your Google Workspace Service Account can read Calendar, Contacts and Gmail, not just Drive.
Frequently asked questions
What is xynehq/xyne?
It is an AI-first search and answer engine for work, described in the README as an open-source alternative to Glean, Gemini and MS Copilot. It connects to applications such as Google Workspace, indexes the data and maps a graph of relationships so users can ask questions and get answers with sources.
How do I install and run Xyne locally?
The README says Xyne runs locally or on a virtual machine with a single docker-compose command, and points to the Quickstart section of the documentation for local use and an AWS EC2 page for cloud deployment. The repository carries start.sh, start-aws-linux.sh and a deployment directory, but the README does not reproduce the compose file itself.
Which LLM can Xyne use, and is my data used for training?
The README states the model layer is agnostic and can point at any LLM or cloud provider, including a local DeepSeek via ollama. It also states that no training happens on your data or prompts and that no telemetry is collected.
Is Xyne still maintained in this repository?
The README states the repository is no longer maintained here and names github.com/juspay/xyne-spaces as the new home. The last push to this repository was on 2026-09-10, and the newest release listed is v3.1.1 from 2025-09-30.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/xynehq-xyne)