healthy-diet-ai-agent: a Bun and TypeScript backend for nutrition chat and RAG diet guidance
Healthy Diet AI Agent is a Bun + TypeScript backend for nutrition chat, food-image analysis, RAG document ingestion, and knowledge-grounded diet guidance.
At a glance
- What is it?
- The repository packages an Express HTTP API and a terminal CLI around a LangChain agent, with SQLite or Supabase as the storage layer. It is a backend service, not a finished nutrition app, and the README leaves several operational questions open.
- Who is it for?
- Adopt it if you want a self-hosted Bun service that already wires an OpenAI-compatible model endpoint to a local SQLite knowledge base and exposes both HTTP and CLI entry points, and if you are willing to read src/server and agent_config.json because the README does not document the API surface. Do not adopt it if you need a supported product, a hosted endpoint, or documented rollback and upgrade procedures; the last push was on 2026-08-16 and there are no retrieved releases.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What healthy-diet-ai-agent actually is, and who it is for
This is a backend, not an app. The README describes a Bun and TypeScript service for nutrition and healthy diet support, with chat, food image analysis, RAG document retrieval, a knowledge graph, and MOHW data synchronization. There is no bundled user interface: the repository ships an Express server entry point at src/index.ts and a CLI entry point at src/cli.ts, plus an openapi.yml file at the top level.
The intended user is a developer who wants to stand up a diet-guidance service on their own machine or server and plug their own frontend into it. The project background section explains that it was originally built to complement PU-Hub/healthy-diet on the API side and archie0732/healthy-diet-web on the frontend, and that the direction was later adjusted so this repository can be deployed independently. Integration with the original stack is retained through Supabase, but it is no longer required.
That history matters for evaluation. The codebase carries two deployment paths, two storage adapters, and a set of modules that exist because of the original pairing. If you are starting fresh, you are paying for that flexibility in configuration surface: STORAGE_BACKEND, SQLITE_DB_PATH, SUPABASE_URL, SUPABASE_SERVICE_KEY, plus a block of RAG worker and MOHW sync variables in .env.example.
How the agent, storage adapters and ingestion worker fit together
The architecture visible in the repository layout is a layered one. src/server/agentRuntime.ts sets up the LangChain and LangGraph agent runtime; src/server/httpRuntime.ts bootstraps the HTTP server; src/serverHandlers.ts holds the router controller handlers. Business logic sits in named modules: knowledgeGraph.ts for extraction and search, knowledgeIngestion.ts for upload, parsing and embedding ingestion, mohwNews.ts for the synchronization task, and ragDocuments.ts for document CRUD and indexer routes.
Storage is abstracted. src/storage/sqlite/ and src/storage/supabase/ are two adapters behind the same interface, selected by STORAGE_BACKEND. That is the cleanest design decision in the repository, because it means the agent code does not branch on the backend. The cost is that any feature added to one adapter has to be mirrored in the other, and the README does not state whether the two are at parity.
Ingestion is asynchronous. .env.example exposes RAG_WORKER_ENABLED, RAG_WORKER_POLL_SECONDS, RAG_WORKER_BATCH_SIZE, RAG_WORKER_MAX_RETRIES, RAG_WORKER_RETRY_BASE_SECONDS, RAG_WORKER_MAX_BACKOFF_SECONDS, RAG_WORKER_PROCESSING_TIMEOUT_SECONDS and RAG_WORKER_STUCK_BATCH_SIZE. The names alone tell you the shape: a polling worker pulls batches, retries with exponential backoff capped at 1800 seconds, times out items after 900 seconds, and has a separate pass for stuck batches. Uploaded files land in knowledge_base/uploads, parsed markdown lands in knowledge_base/ingested_markdown.
There is also a version-aware and policy-aware RAG evaluation engine, with a frozen suite under experiments/version_aware_rag/. The README frames its purpose as resolving multi-version dietary guideline temporal conflicts. That is the most interesting claim in the project, and it is also the least documented here: the README does not describe the retrieval rules or how conflicts are resolved at query time, so you would need to read the experiments directory to judge it.
Running healthy-diet-ai-agent with Docker Compose and SQLite
The repository ships a compose.yml that defaults to standalone SQLite mode. It pulls ghcr.io/archie0732/healthy-diet-ai-agent:main, sets PORT to 8001 and STORAGE_BACKEND to sqlite, mounts ./data, ./knowledge_base and ./users_images, and adds host.docker.internal:host-gateway so the container can reach a model server running on the host. The Dockerfile builds from oven/bun:1.2.12, installs production dependencies with bun install --frozen-lockfile --production, and declares a healthcheck that fetches /ping.
Start by copying the example environment file, because compose.yml reads .env through env_file and will fail without it:
cp .env.example .envThen set AI_API_URL to your OpenAI-compatible endpoint. The example points at a host-local server on port 8080:
AI_API_URL=http://host.docker.internal:8080/v1/Bring the stack up and check the health endpoint the Dockerfile itself uses:
docker compose up -d
curl http://localhost:8001/pingThe README does not document the response body of /ping, only that the Dockerfile treats a non-ok HTTP status as unhealthy. If the container restarts in a loop, the first thing to check is whether your model endpoint is reachable from inside the container, since the agent runtime has nothing to talk to otherwise.
For local development without Docker, package.json defines the scripts. bun run dev and bun run start both execute bun ./src, which is the HTTP server. The CLI is separate:
bun run cliThe CLI reads CLI_USER_ID and CLI_THREAD_ID from the environment, defaulting to local-user and local-thread in .env.example. Those two values are what separate conversation threads, so running the CLI twice with the same CLI_THREAD_ID continues the same thread.
Model routing, Gemini fallback and the agent_config.json split
Configuration is deliberately split in two. .env.example states that repo-level role, prompt, MOHW and RAG defaults live in agent_config.json, and that environment values should be used only for deployment-specific overrides. That is a sensible separation, and it means a fork that only changes behaviour does not need to touch the environment at all.
The model routing has a wrinkle worth reading carefully. The comment in .env.example says the current repo code reads GEMINI_AI_API first, and that GEMINI_API_KEY is only a fallback in src/server/modelRouting.ts. The example file also sets GOOGLE_CHAT_MODEL to gemma-4-31b-it and GOOGLE_BASE_URL to the Google Generative Language OpenAI-compatibility endpoint. If you set GEMINI_API_KEY expecting it to take effect, the README gives you no reason to think it will. Set GEMINI_AI_API.
There is a second routing detail that the README does not resolve: it lists @langchain/tavily as a dependency, and the agent framework list includes DeepAgents alongside LangChain and LangGraph. Neither Tavily search nor the DeepAgents workflow is described in the feature list, so it is unclear from the README alone whether they are active in the default agent configuration or present for optional use. Treat that as something to confirm in agent_config.json and src/server/agentRuntime.ts before you assume web search is part of the loop.
On the positive side, routing through an OpenAI-compatible interface means the service is not tied to one vendor. Any server that speaks that API shape can sit behind AI_API_URL, and the Gemini path is an additional option rather than a requirement.
Where healthy-diet-ai-agent is the wrong tool
The most concrete limitation is that this is not a deployable product. There is no homepage, no retrieved release, and no published API reference beyond an openapi.yml file at the repository root. The README's Deployment section is truncated in the repository listing, so the documented install path is essentially the compose file and the package.json scripts. If you need a supported service with a versioned changelog and a compatibility promise, this is not it. The last push was on 2026-08-16.
Second, nutrition guidance is a domain where wrong answers carry real cost. The repository contains knowledge_base/NUTRITION_RULES.md, described in the layout as ground-truth guidelines for dietary analysis, and the agent is meant to be grounded in ingested documents. The README does not describe what happens when retrieval returns nothing relevant, or whether the agent refuses to answer outside the knowledge base. That gap matters more here than in a general-purpose chat backend.
Third, the MOHW synchronization pipeline is jurisdiction-specific. The README identifies MOHW as the Ministry of Health and Welfare and describes the pipeline as importing public clarification and reference content. If you are not operating in that regulatory context, that module and its three sync environment variables are dead weight you will need to disable or ignore.
Finally, the food image analysis path implies image handling through sharp and a multimodal model. The README lists the workflow as a feature but gives no accuracy figures, no supported image formats, and no size limits. Do not plan around it until you have read the handler.
How it differs from wiring LangChain to a vector store yourself
The obvious alternative is assembling the same thing from LangChain and a vector database directly, which is what this project does internally with @langchain/core, @langchain/langgraph and @langchain/openai. The difference is not the framework, it is what comes pre-wired: a storage abstraction with two adapters, an ingestion worker with retry and stuck-batch handling, a document CRUD route set, a knowledge graph module, a PDF-to-Markdown conversion script at src/rag_clean/pdf_to_clean_markdown.py, and a CLI.
Building that yourself is maybe a few days of work, and you would get exactly the pieces you need. Adopting this gives you the pieces someone else decided were needed, including the MOHW sync and the version-aware RAG experiment suite. The trade is control over scope versus time to first query. If your retrieval needs are simple, a plain LangChain script plus a vector store is less code to maintain than a service with two storage backends.
A second alternative is using a hosted nutrition or health assistant API. That removes the ingestion, embedding and model-routing work entirely. It also removes the ability to keep documents on your own disk, which is the main reason to pick a self-hosted backend like this one. The choice is really about data locality, not features.
Licence, upgrade cost and what the repository does not promise
The licence is MIT, declared in both package.json and the LICENSE file, and the README carries an MIT badge. In practical terms that permits commercial use and modification with attribution and no warranty. This is not legal advice; read the LICENSE file if the distinction matters to you.
Upgrade cost is the weak point. There are no retrieved releases, so there is no version tag to pin to and no release notes to read before upgrading. compose.yml sets pull_policy: always on the main tag, which means docker compose up will fetch whatever is currently on main rather than a fixed build. For a service holding your knowledge base in a mounted volume, that is a meaningful operational risk: an upstream change to the SQLite schema under docs/sqlite/ could land without a migration step you were warned about. If you deploy this, pin IMAGE_NAME to a digest or a specific tag rather than accepting the default.
Dependencies are another cost. The stack includes LangChain 0.3.x alongside @langchain/core 1.x and @langchain/langgraph 1.x, plus deepagents, sharp, pdf-parse and mammoth. That is a wide surface for a backend, and the version spread between the langchain package and the @langchain scoped packages is the kind of thing that produces resolution conflicts during upgrades. Run bun install --frozen-lockfile as the Dockerfile does rather than letting Bun resolve fresh versions.
Editorial conclusion
Adopt it if you want a self-hosted Bun service that already wires an OpenAI-compatible model endpoint to a local SQLite knowledge base and exposes both HTTP and CLI entry points, and if you are willing to read src/server and agent_config.json because the README does not document the API surface. Do not adopt it if you need a supported product, a hosted endpoint, or documented rollback and upgrade procedures; the last push was on 2026-08-16 and there are no retrieved releases. Verify first that your model endpoint answers at the /v1/ path you put in AI_API_URL, that the container can reach it through host.docker.internal, and that GET /ping returns ok before you point any client at port 8001.
Frequently asked questions
Is healthy-diet-ai-agent an AI nutrition tool I can use directly?
It is a backend service, not an end-user application. The README describes a Bun and TypeScript backend with an HTTP API and a terminal CLI, and the project background notes that a separate frontend repository, archie0732/healthy-diet-web, was built alongside it.
Is healthy-diet-ai-agent free to run?
The code is MIT licensed, so there is no licence fee. You still supply an OpenAI-compatible model endpoint through AI_API_URL, and any cost from that provider is yours; the README does not describe a hosted free tier.
Does healthy-diet-ai-agent replace a nutritionist?
The README makes no such claim and does not describe clinical validation. It presents the service as grounded in ingested documents, including knowledge_base/NUTRITION_RULES.md, and does not state what the agent does when retrieval returns nothing relevant.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/archie0732-healthy-diet-ai-agent)