CLI tool
skaldlabs/skald avatar
skaldlabs/skald

Skald: a self-hosted RAG API you run in your own infrastructure

Context layer platform in your infrastructure

565 stars38 forksTypeScriptNOASSERTION

At a glance

What is it?
Skald is a context layer platform that wraps document ingestion and retrieval behind an HTTP API, with SDKs for Python, Node, Go, Ruby, PHP, C# and MCP. This review covers how it is deployed, what the defaults assume, and where the trade-offs sit.
Who is it for?
Adopt Skald if you want retrieval and chat behind a single API without building the ingestion pipeline yourself, and if you are comfortable running Postgres with pgvector plus RabbitMQ alongside it. Do not adopt it if you need a fully offline deployment out of the box, or if a single-language stack with no message broker is a hard requirement.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
No. The owners have archived the repository on GitHub, so it is read-only and no longer receives changes.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Skald targets, and who it is actually for

Retrieval augmented generation has a boring middle. Parsing documents, chunking them, generating summaries and tags, embedding them, storing vectors, rewriting queries, running the search, calling a model, keeping chat history and attaching source references. None of it is intellectually hard in isolation. All of it is tedious to assemble and easy to get subtly wrong.

Skald's pitch is that this middle should be an API call. The README describes the ingestion phase as covering document parsing, chunking strategy, summaries, tagging, embedding generation and vector storage, and the retrieval phase as covering query rewriting, vector search, LLM chat, chat history and source references. The intended user is a product team that wants a working RAG endpoint this week and a configuration surface later, not a research group building a retrieval system from primitives.

The self-hosted framing matters. The repository ships docker-compose files and an embedding-service directory, and the README states you can deploy without any third-party dependencies if you host your own inference server and use the local embeddings service in the `local` docker compose profile. That is a meaningful distinction from API-only RAG products, where the corpus leaves your network by design.

How the pieces fit: Postgres, RabbitMQ, an API and a separate UI

The repository layout tells you more than the marketing copy. There is a backend directory, a frontend directory, an embedding-service directory, an ee directory, and three compose files: docker-compose.yml, docker-compose.selfhosted.yml and docker-compose.ee.yml. The root package.json is a private Vite and React application, not a published library, which means the top-level npm project is the platform UI rather than the server.

The default docker-compose.yml defines three services. A db service runs `pgvector/pgvector:pg17` with a POSTGRES_DB of skald, mounts ./docker-init into the entrypoint directory, and exposes 5432. A rabbitmq service runs `rabbitmq:3.13-management-alpine`, exposes 5672 for AMQP and 15672 for the management UI, and uses the default guest credentials. The api service builds from ./backend, runs `sh /app/start.sh`, and takes its configuration from a .env file plus environment variables including EXPRESS_SERVER_PORT=8080, DB_HOST, LLM_PROVIDER, EMBEDDING_PROVIDER and RERANKING_PROVIDER.

That broker is the architectural detail worth pausing on. Ingestion is asynchronous, so RabbitMQ sits between the API accepting a document and the worker that parses, embeds and stores it. It is a reasonable design for slow embedding calls, but it also means a self-hosted Skald is a three-process system with a queue to monitor, not a single binary. The README does not document what happens to in-flight ingestion jobs when the broker restarts.

Installing Skald locally with docker-compose

The README gives a four-command quick start. It clones the repository, writes an OpenAI key into a .env file, and brings the stack up with docker-compose.

bash
git clone https://github.com/skaldlabs/skald
cd skald
echo "OPENAI_API_KEY=<your_key>" > .env
docker-compose up

The README states this is enough for a local run. Because the api service reads .env through env_file and the compose file defaults LLM_PROVIDER and EMBEDDING_PROVIDER to openai, that single key is what the stack expects unless you override the providers.

For a production self-hosted deployment the README points to the self-hosting documentation rather than the quick start, and the repository provides docker-compose.selfhosted.yml for that path. The README claims a fully-featured production deploy with SSL can be live in less than an hour. That is a vendor estimate, not something this review can confirm.

Once the stack is up, the README's Node example is the shortest path to a first real use. It creates a memo and then asks a question against it.

js
import { Skald } from '@skald-labs/skald-node'

const skald = new Skald('your-api-key-here')

await skald.createMemo({
    title: 'Meeting Notes',
    content: 'Full content of the memo...',
})

const chatRes = await skald.chat({
    query: 'What were the main points discussed in the Q1 meeting?',
    rag_config: { reranking: { enabled: true } },
})

console.log(chatRes.response)

The README shows this snippet printing `chatRes.response`. Note that the snippet passes an API key to the constructor, while the quick start only sets OPENAI_API_KEY. The .env.example includes a DISABLE_AUTH flag, which suggests local runs can skip authentication, but the README does not connect those two facts, so treat the API key value as something you must resolve from the self-hosting docs.

Provider configuration is where the defaults stop being plug-and-play

The .env.example is the most informative file in the repository, and it is also where the easy story gets complicated. It lists three embedding providers (voyage, openai, local) and four LLM providers (openai, anthropic, groq, local), each with its own key and model variables: VOYAGE_EMBEDDING_MODEL, OPENAI_EMBEDDING_MODEL, OPENAI_MODEL, ANTHROPIC_MODEL, GEMINI_MODEL, GROQ_MODEL, and an EMBEDDING_SERVICE_URL pointing at http://localhost:8001 for the local embedding service.

The compose file defaults RERANKING_PROVIDER to local, which is a different setting from the embedding and LLM providers. Nothing in the README explains what the local reranker runs on or how much memory it needs. That is a gap: reranking is one of the features the README advertises as tunable, and the default value points at a component the README does not describe.

The file also carries configuration that has nothing to do with RAG: Google OAuth credentials, a Resend key for email, Stripe secret, publishable and webhook keys, and EMAIL_VERIFICATION_ENABLED. These belong to the hosted platform rather than to a pure retrieval service. A self-hoster who only wants an API endpoint is looking at a configuration surface built for a multi-tenant product, and it is not obvious from the README which of those variables are safe to leave blank.

Where Skald is the wrong choice

The strongest limitation is spelled out by the project itself. Running without any third-party services is described in the README as advanced usage that requires hosting your own LLM inference server, and the local embedding path requires the separate embedding-service container. If your requirement is a genuinely offline deployment with no external calls, Skald can do it, but you are signing up to operate an inference stack as well as a RAG stack.

A second limitation is operational weight. A minimal Skald deployment is an Express API, Postgres with pgvector, RabbitMQ, and optionally a Python embedding service. Teams that want a single container with an embedded vector store will find that footprint disproportionate. Teams already running Postgres and a broker will not.

Third, the README does not document rollback, re-ingestion, or what happens when you change EMBEDDING_PROVIDER or the embedding model on an existing corpus. Switching from voyage-3-large to text-embedding-3-large changes the vector space, and the README is silent on whether Skald re-embeds stored documents or leaves them inconsistent. That is a question to settle before you ingest anything you care about.

Finally, chunking is listed in the README as configurable "soon", which means the chunking strategy is currently fixed. If your documents need a specific structure-aware split, Skald's defaults are what you get.

How it compares to assembling the same pipeline yourself

The obvious alternative is not another product but a hand-built stack: a parsing library, a chunker, an embedding client, pgvector or another vector store, and a chat loop. That approach gives you exact control over chunk boundaries, metadata and prompt construction, and it adds no message broker or API layer you did not write. The cost is that you own query rewriting, chat history persistence and source attribution, which is precisely the surface Skald exposes as configuration.

A second alternative is the hosted RAG API category, where you send documents to a vendor endpoint and get answers back. The difference in approach is directional: hosted services keep the pipeline outside your network and remove the operational burden, while Skald's README explicitly positions self-hosting as the point, with a cloud option at useskald.com for those who want the managed path. Choosing Skald self-hosted means choosing the operational burden on purpose, usually because the corpus cannot leave your infrastructure.

A third option is a framework library that runs in your own process rather than behind an HTTP API. That keeps everything in one language and one deployment unit, at the cost of reimplementing the API, auth and multi-client story that Skald's SDKs cover across Python, Node, Go, Ruby, PHP, C# and MCP.

Licence, maintenance and the cost of staying current

The README states the repository is MIT licensed except for the ee directory, which has its own licence file. It also points to a separate skald-foss repository containing the same code with the ee directory removed, and states that skald-foss is what runs the Cloud offering today, with the Enterprise Edition aimed at on-prem deployments. For anyone who needs unambiguously permissive code, the README's own recommendation is to use skald-foss rather than the main repository. Read both licence files before you decide; this is a description of what the README says, not legal advice.

On maintenance, the last push to the default branch was on 2026-05-31. The README advertises SDK versions across six languages plus MCP and a CLI, which means upgrades are not a single-event concern: an API change can require coordinated bumps in whichever SDKs you consume. The README does not describe a versioning or deprecation policy for the HTTP API, so pin your SDK versions and read the release notes for each one independently.

The upgrade cost that is easiest to underestimate is the embedding model. Because the model is set through environment variables and the README does not document a re-embedding path, treat the choice of EMBEDDING_PROVIDER and its model as a decision with a migration attached.

Editorial conclusion

Adopt Skald if you want retrieval and chat behind a single API without building the ingestion pipeline yourself, and if you are comfortable running Postgres with pgvector plus RabbitMQ alongside it. Do not adopt it if you need a fully offline deployment out of the box, or if a single-language stack with no message broker is a hard requirement. Before committing, read the self-hosting docs, check which directory you are cloning (the README points at both skaldlabs/skald and skaldlabs/skald-foss), confirm the licence terms for the ee directory, and verify that your chosen embedding and LLM providers are among the ones the environment file lists.

Frequently asked questions

What does Skald do?

Skald is a context layer platform that handles document ingestion (parsing, chunking, summaries, tagging, embeddings, vector storage) and retrieval (query rewriting, vector search, LLM chat, chat history, source references) behind an API. The repository ships docker-compose files for self-hosting and SDKs for several languages.

What is Skald?

According to the README, Skald is an open source production RAG system that runs in your own infrastructure, exposed through a plug-and-play API with configurable vector search parameters, reranking, models and query rewriting.

How do I install Skald?

The README's quick start clones the repository, writes an OPENAI_API_KEY into a .env file, and runs docker-compose up. For production, the README points to the self-hosting docs and the docker-compose.selfhosted.yml file.

Can Skald run without OpenAI or any other third-party service?

The README states you can deploy without any third-party dependencies, including OpenAI, but that this requires hosting your own LLM inference server and using the local embeddings service provided in the local docker compose profile. The README calls this advanced usage.

Which languages have a Skald SDK?

The README lists Python, Node, Ruby, Go, PHP and C# SDKs plus an MCP integration and a CLI, each with its own version badge. The README says to open an issue if your language is missing.

Is Skald open source?

The README states the repository is MIT licensed except for the ee directory, which has its own licence. It also points to the skald-foss repository, which contains the same code with the ee directory removed.

Official sources

  1. Issues
  2. Project website
  3. README
  4. skaldlabs/skald on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/skaldlabs-skald.svg)](https://hysenlabs.com/projects/skaldlabs-skald)