Model or dataset
llm-tools/embedJs avatar
llm-tools/embedJs

EmbedJs: a Node.js RAG framework for wiring your own data into an LLM

A NodeJS RAG framework to easily work with LLMs and embeddings

601 stars76 forksTypeScriptApache-2.0

At a glance

What is it?
EmbedJs splits documents into chunks, embeds them and stores them in a vector database so a Node.js app can answer questions from its own content. The npm package is @llm-tools/embedjs, and the repository is a monorepo of swappable loaders, models and databases.
Who is it for?
Adopt EmbedJs if you are building a Node.js service and want the chunk, embed, store and query loop assembled for you, with loaders and vector stores you can swap. Do not adopt it if you need a hosted pipeline, a Python stack, or a stable API surface: the version is 0.1.31 and the last push was on 2026-06-26.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 95 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What EmbedJs is for, and who ends up using it

An LLM answers from its training data. If you want it to answer from your Confluence space, your markdown files or a folder of images, you have to retrieve the relevant passages yourself and put them in the prompt. EmbedJs is a Node.js framework for that retrieval step. The README describes it as a framework "for personalizing LLM responses" and as a toolkit for building RAG and LLM applications in Node.js.

The audience is a JavaScript or TypeScript developer who already has an LLM API key and a corpus, and who does not want to hand-roll chunking, embedding calls and vector search. The repository is a TypeScript monorepo with workspaces named core/*, databases/*, loaders/* and models/*, so the intended shape is a small core plus pluggable pieces. The examples directory confirms the intended entry points: simple, dynamic, markdown, image, libsql, pinecone and confluence.

It is not a hosted product and not a chatbot UI. It is the library layer underneath one.

The chunk, embed, store, query loop

The README states the pipeline in one sentence: it "segments data into manageable chunks, generates relevant embeddings, and stores them in a vector database for optimized retrieval." That is the whole architecture. A loader reads a source and produces text. The core splits that text into chunks. A model turns each chunk into a vector. A database adapter writes the vector and its payload. At query time the same model embeds the question, the database returns the nearest chunks, and those chunks are handed to the LLM as context.

The monorepo layout is the interesting part. loaders/ holds the source adapters, models/ holds the embedding and LLM providers, databases/ holds the vector stores. The topics list names OpenAI, Claude, Cohere, HuggingFace, Mistral, Ollama, Vertex AI, Pinecone and general vector databases, so the provider set is broad by design. Swapping a model or a vector store should mean changing which package you depend on rather than rewriting the loop.

Because the pieces are separate npm packages under the @llm-tools scope, the core package alone will not get you a working pipeline. You pick a loader for your data, a model for embeddings and a database for storage.

Installing @llm-tools/embedjs and running a first query

The README points to the npm registry badge for the package name and to the quickstart page for setup steps. The package is published as @llm-tools/embedjs. The repository does not include a full install transcript in the README, so treat the quickstart page as the authoritative source for the current package list and configuration keys.

Start by adding the core package to a project that already has an LLM provider credential available as an environment variable. The repository uses ESM: package.json sets "type": "module", so import syntax is the expected form.

bash
npm install @llm-tools/embedjs

The repository ships runnable samples under examples/. The simple example is the smallest one, and the repository root exposes a script that runs them, so the intended way to see the loop work is to run an example rather than to write one from scratch.

bash
npm run example

If you prefer to follow the documented path, the quickstart is at llm-tools.mintlify.app/get-started/quickstart, and the examples index is at llm-tools.mintlify.app/examples. The README also links a supported data types page under components/data-sources/overview. Check that page before writing a loader, because the list of loaders is the part of the framework most likely to have changed since the version you are reading about.

Where the framework gets in your way

The version number is the first thing to weigh. The most recent release listed is v0.1.31, published on 2025-11-14. A 0.1.x line means the API is not promised to be stable, and the gap between v0.1.29 in June 2025 and v0.1.30 in November 2025 shows that releases arrive in bursts rather than on a schedule. The last push to the repository was on 2026-06-26, so the code has moved since the latest release. Pin an exact version and read the changelog before upgrading.

The second constraint is the split packaging. If your deployment target is a single bundle, you are assembling several @llm-tools packages yourself, and each one carries its own provider SDK. That is more dependency surface than a single-package framework, and it is the cost of the pluggable design.

The third is scope. The README describes retrieval and chat over your own data. It does not describe evaluation harnesses, reranking, query rewriting or answer-quality metrics. If your problem is that retrieval returns the wrong chunks rather than that retrieval does not exist, EmbedJs gives you the loop but not the diagnosis. You will be measuring relevance yourself.

Finally, the documentation is a separate site, not the README. The README is short and mostly links out. Anything you need beyond the one-paragraph description lives at llm-tools.mintlify.app, so an offline or air-gapped workflow is harder than it looks.

Compared with LangChain.js and LlamaIndex.TS

The closest alternatives in the same language are LangChain.js and LlamaIndex.TS. All three sit in the same slot: a JavaScript layer that connects data sources to LLMs through embeddings.

The difference is how much of the application they claim. LangChain.js is a general orchestration library: chains, agents, tools, memory and output parsers, with RAG as one supported pattern among many. LlamaIndex.TS centres on indexing and query engines, with a strong emphasis on how documents are structured and queried. EmbedJs is narrower and more literal about the pipeline: loaders, models, databases, and the chunk-embed-store-retrieve loop. The README does not promise agents or tool calling.

That narrowness is the argument for it. If you want a RAG endpoint inside an existing Node service and you do not want a general orchestration framework in your dependency tree, the smaller surface is easier to reason about. If you expect to add agents, tool use or multi-step planning later, you will likely outgrow it and should start with a broader library instead. Choosing EmbedJs is a bet that retrieval is the whole job.

Licence, maintenance and the cost of upgrading

The repository is licensed Apache-2.0, and the npm package reports the same licence. Apache-2.0 is a permissive licence with an explicit patent grant, which matters if your company has a review process for dependencies. It is not a copyleft licence, so it does not oblige you to publish your application source. This is a description of the licence text, not legal advice; your own counsel decides how it applies to your distribution model.

On maintenance, the facts are these: the repository is not archived, the last push was on 2026-06-26, and the latest release in the list is v0.1.31 from 2025-11-14. That is a repository with recent commits and a release cadence that lags behind them, which is normal for a monorepo that publishes several packages.

The upgrade cost is structural. Because the framework is split into core, loaders, models and databases workspaces, a breaking change in a provider SDK can force a version bump in the corresponding @llm-tools package without touching the others. Before upgrading, check which of your installed @llm-tools packages changed and whether the loader or database you rely on is among them. Pinning exact versions in package.json is the cheap insurance.

Editorial conclusion

Adopt EmbedJs if you are building a Node.js service and want the chunk, embed, store and query loop assembled for you, with loaders and vector stores you can swap. Do not adopt it if you need a hosted pipeline, a Python stack, or a stable API surface: the version is 0.1.31 and the last push was on 2026-06-26. Before writing code against it, read the quickstart at llm-tools.mintlify.app/get-started/quickstart and confirm which loader and database packages your use case needs, because the framework is published as separate workspace packages rather than one bundle.

Frequently asked questions

What is EmbedJs used for?

It is a Node.js framework for personalizing LLM responses from your own data. It segments data into chunks, generates embeddings and stores them in a vector database so an application can retrieve context and answer questions or hold a chat over that corpus.

How do I install EmbedJs?

The package is published on npm as @llm-tools/embedjs, and the README links a quickstart page at llm-tools.mintlify.app/get-started/quickstart for the current setup steps. The repository is ESM, since package.json sets "type": "module".

Which vector databases and model providers does EmbedJs support?

The repository is organised into databases/*, loaders/* and models/* workspaces, and the GitHub topics name OpenAI, Claude, Cohere, HuggingFace, Mistral, Ollama, Vertex AI and Pinecone. The supported data types page under components/data-sources/overview lists the loaders.

Is EmbedJs production ready?

The latest release listed is v0.1.31, published on 2025-11-14, so the API is still on a 0.1.x line and is not promised to be stable. The repository is not archived and the last push was on 2026-06-26.

Is EmbedJs free to use in a commercial product?

The repository and the npm package are both licensed Apache-2.0, a permissive licence that does not require you to publish your application source. How that applies to your distribution is a question for your own legal review.

Official sources

  1. License: Apache-2.0
  2. llm-tools/embedJs on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/llm-tools-embedjs.svg)](https://hysenlabs.com/projects/llm-tools-embedjs)