Model or dataset
xr843/fojin avatar
xr843/fojin

FoJin: a self-hosted Buddhist canon with cited AI answers

Buddhist Digital Text Platform — 10,500+ texts, 613 sources, trilingual cross-canon, AI Q&A (RAG), knowledge graph, full-text search

345 stars64 forksPythonApache-2.0

At a glance

What is it?
FoJin aggregates 612 sources into one searchable corpus and answers questions with clickable citations. Here is how the Docker stack is put together, what the RAG pipeline promises, and where the documentation stops.
Who is it for?
Adopt FoJin if you need a self-hosted, citation-first Buddhist corpus and can supply an OpenAI-compatible LLM endpoint plus the Postgres, Elasticsearch and Redis services the compose file expects. Skip it if you want a small reading app, or if you cannot run a multi-container stack with 4 GB of Postgres memory and a 256 MB shared-memory segment.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem FoJin targets: one question, hundreds of silos

Buddhist primary texts live in separate databases with separate interfaces. The README names CBETA, SuttaCentral, BDRC, SAT, 84000 and GRETIL as examples, and describes the cost plainly: you spend more time locating the right passage than reading it. FoJin's answer is aggregation plus retrieval. It reports 612 data sources folded into one index, 10,500+ texts, and 19,000+ volumes of full content in Classical Chinese, Pali, Tibetan and Sanskrit. The audience is narrow on purpose: Buddhist studies researchers, digital humanities engineers, and developers who want the corpus exposed as an API rather than a website. The README frames FoJin as "open, cross-canon, verifiable Buddhist knowledge infrastructure", and the API surface backs that framing. Every passage is meant to carry a stable URN such as fojin:cbeta/T0001.1, so a citation resolves to a specific text rather than a search result. That design choice is what separates FoJin from a full-text search box: the unit of retrieval is addressable and citable. If you only want to read sutras, the aggregation is still useful, but the URN layer and the verification endpoint add machinery you will not exercise.

How the retrieval pipeline is wired

The README describes a Retrieval-Augmented Generation loop over 670K+ embedded passages, with optional cross-encoder reranking and root-sutra recall. The knowledge graph holds 110K+ entities and 27,900+ relations, and the semantic similarity feature is stated to run on pgvector with HNSW indexes. So there are at least three retrieval paths sharing one corpus: vector similarity in Postgres, lexical search in Elasticsearch, and graph traversal for entities and lineage. A separate Research Assistant at /research is described as planning across corpus, dictionaries and knowledge graph before synthesising an answer. The trust layer is the part worth reading closely. FoJin applies a deterministic citation whitelist, downgrades answers whose quotes are not verbatim, and assigns a per-answer trust state; the README claims roughly 98% of citing answers are served-trustworthy at temperature 0. Treat that figure as a project claim, not an independent measurement. The mechanism behind it is checkable in principle, because the same verbatim check is exposed publicly: /api/verify/quote takes a sentence and a citation and answers two separate questions, whether the quote is real and whether the citation is right, returning exact, near-miss or absent plus a character-level diff. Exposing the guard as an endpoint is the most interesting decision in the project. It lets a third party test the same property the product advertises.

Installing FoJin with Docker Compose

The repository ships docker-compose.yml, .env.example, deploy.sh and deploy/. The compose file defines at least postgres (image pgvector/pgvector:pg15), elasticsearch (built from ./elasticsearch) and Redis, with backend and frontend port mappings. Start by copying the example environment file. Note that POSTGRES_PASSWORD is required by the compose file, so an unset value fails fast rather than starting with a default.

bash
cp .env.example .env
python -c "import secrets; print(secrets.token_urlsafe(32))"

The second command generates a value you can paste into JWT_SECRET_KEY. The example file also lists API_KEY_ENCRYPTION_KEY, with a Fernet generation one-liner in its comment, and states that in development an ephemeral key is generated per boot, so stored ciphertexts do not survive a restart. For a real deployment, set FOJIN_ENV and provide strong secrets: the comments say that when FOJIN_ENV is unset the app treats the environment as production and refuses to start without them.

bash
# FOJIN_ENV=production
# JWT_SECRET_KEY=<generated above>
# API_KEY_ENCRYPTION_KEY=<Fernet key>
# LLM_API_KEY=sk-your-llm-api-key
# LLM_API_URL=https://api.deepseek.com/v1

The LLM settings are OpenAI-compatible, so the Q&A features need a reachable endpoint and key before they will answer anything. Then bring the stack up. The compose comments warn that changing shm_size for Postgres needs an explicit restart of that service, because deploy.sh does not recreate postgres on a compose-only change.

bash
docker compose up -d postgres
docker compose up -d

The README points at https://fojin.app for a live demo and https://fojin.app/docs for API documentation, so you can compare your local build against the hosted one. For a first real use, query the verification endpoint rather than the chat: it needs no model call and tells you immediately whether your corpus and citation index are working.

The resource envelope is not small

Read the compose file before promising anyone a quick install. Postgres is capped at 4 GB of memory and one CPU, with a comment explaining that 1 GB was chronically OOM-killed and that raising it was necessary to stop crash-recovery cycles. The same block sets shm_size to 256 MB because Docker's 64 MB default caused parallel VACUUM on a table around 800 MB to fail with a shared-memory error. Elasticsearch is capped at 1536 MB and 1.5 CPUs. That is a floor for a single-node evaluation, before you have loaded the full corpus, and the comments imply the authors hit these limits in production rather than guessing them. A second constraint is data provenance. FoJin aggregates 612 sources, and the README does not document which sources ship with the repository versus which are fetched or licensed separately. The datasets/ directory exists, but the README does not describe its contents or ingestion steps. If your question depends on a specific canon edition, verify it is present in your build before trusting an answer. Finally, the AI features are only as good as the configured model. The README does not document behaviour when the LLM endpoint is unreachable, so plan for the Q&A surface to be unavailable independently of search.

fojin-mcp and the case for calling FoJin instead of rebuilding it

The realistic alternative is not another Buddhist platform. It is building your own pipeline over CBETA or SuttaCentral dumps, or simply using those projects' own search interfaces. The difference is architectural. CBETA and SuttaCentral are primary sources with their own reading and search tools; neither is described here as offering a citation-verification endpoint or a cross-canon URN scheme. FoJin's bet is that the aggregation plus the verification layer is the product, and the README supports that by publishing fojin-mcp on PyPI, hosted at mcp.fojin.ai as an anonymous service with no key required, exposing eight read-only, URN-addressable tools to any MCP client. If you already run a RAG stack, the MCP server is the cheaper integration path: you get cited passages without operating Postgres, Elasticsearch and Redis yourself. If you need offline operation or your institution forbids external calls, the self-hosted compose stack is the only option, and then you own the resource envelope described above. The honest trade-off is that a custom pipeline gives you exact control over which editions you ingest, while FoJin gives you breadth and a citation contract you did not have to design.

Maintenance, licensing and upgrade cost

The repository is not archived, and the last push was on 2026-09-10. Releases are tagged frequently, with v1.0.0 on 2026-03-23, v4.0.0 on 2026-03-19 and v3.4.0 on 2026-03-13, so the version numbers are not sequential in date order; do not assume v4.0.0 supersedes v1.0.0 without reading CHANGELOG.md. Upgrade cost is dominated by the database. The compose comments show that memory and shared-memory settings were tuned in response to real failures, which means a version bump can require revisiting those values, and the note about deploy.sh not recreating postgres on a compose-only change is exactly the kind of detail that turns a routine pull into a manual step. The repository includes fojin-backup.sh and fojin-eval-regression.sh, so backup and an evaluation regression run are part of the intended workflow; the README does not document rollback. Licence is Apache-2.0, with a NOTICE file present. Apache-2.0 permits commercial use and modification and includes a patent grant, but it also requires preserving notices and stating changes. The corpus itself is a separate question: the licence covers the code, and the README does not state the licensing status of the aggregated source texts, which come from many upstream projects. Check each source's terms before redistributing content.

Editorial conclusion

Adopt FoJin if you need a self-hosted, citation-first Buddhist corpus and can supply an OpenAI-compatible LLM endpoint plus the Postgres, Elasticsearch and Redis services the compose file expects. Skip it if you want a small reading app, or if you cannot run a multi-container stack with 4 GB of Postgres memory and a 256 MB shared-memory segment. Before committing, check the deploy and backup scripts, confirm which of the 612 sources your build actually ingests, and read ARCHITECTURE.md and DECISIONS.md for the parts the README leaves out.

Frequently asked questions

What is FoJin and what does it do?

FoJin is a Buddhist digital text platform that aggregates 612 data sources into one searchable corpus of 10,500+ texts. Its AI assistant, XiaoJin, answers plain-language questions with clickable citations that open the source passage.

How do I install and run FoJin locally?

The repository provides docker-compose.yml and .env.example. Copy the example file, set POSTGRES_PASSWORD, JWT_SECRET_KEY and API_KEY_ENCRYPTION_KEY, then run docker compose up -d; the compose file defines postgres, elasticsearch and Redis services.

Does FoJin require an LLM API key?

The .env.example file lists LLM_API_KEY and LLM_API_URL with an OpenAI-compatible endpoint as the default. The AI Q&A features depend on that endpoint being reachable, while search and reading features do not.

What licence is FoJin released under?

The repository is licensed under Apache-2.0 and includes a NOTICE file. The README does not state the licensing status of the aggregated source texts, which come from many upstream projects.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. xr843/fojin on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/xr843-fojin.svg)](https://hysenlabs.com/projects/xr843-fojin)