Library / SDK
agentset-ai/agentset avatar
agentset-ai/agentset

Agentset: An Open-Source RAG Platform with Built-in Citations and MCP

The open-source RAG platform: built-in citations, deep research, 22+ file formats, partitions, MCP server, and more.

2,099 stars187 forksTypeScriptMIT

At a glance

What is it?
Agentset is an MIT-licensed TypeScript platform for building and deploying retrieval-augmented generation applications, with a chat playground, built-in citations, multi-tenancy, and an MCP server. Self-hosting requires a stack of at least six external services, making the managed cloud tier at agentset.ai the faster starting point for most teams.
Who is it for?
Teams building production RAG applications who want an open-source codebase they can inspect and extend are the right audience. The cloud tier at app.agentset.ai provides a free starting point without configuring any external services, and the self-host path documents the full .env.example for teams that need data residency control.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The Production RAG Gap Agentset Fills

Building a retrieval-augmented generation application from scratch requires assembling ingestion, chunking, embedding, vector indexing, retrieval, and a serving layer, then adding evaluation tooling, multi-tenancy support, and a way to show citations to users. Most open-source projects address one layer at a time.

Agentset packages all of these into a single platform. The README describes it as end-to-end tooling for building, evaluating, and shipping production-ready RAG and agentic applications. The feature list covers ingestion, chunking, embeddings, retrieval, a chat playground with message editing and per-message citations, production hosting with preview links and custom domains, typed SDKs, an OpenAPI spec, and built-in multi-tenancy.

The intended users are TypeScript developers and engineering teams who need to ship a document question-answering or agentic research product and want a foundation that handles infrastructure concerns rather than building each component individually.

The .env.example file in the repository gives a concrete picture of what production deployment looks like. It lists variables for PostgreSQL, two vector store options (Pinecone and Turbopuffer), Azure-hosted embeddings, Cohere for re-ranking, a partitioning API for document ingestion, S3-compatible storage for file uploads, Upstash Redis, Stripe for billing, Discord webhooks for alerts, and Tinybird for analytics. This is the full dependency surface of a self-hosted Agentset installation.

Architecture: A Bun Monorepo with Turbo Build Orchestration

The repository is a monorepo using Bun as the package manager and Turbo for build orchestration. The top-level package.json names the project app.agentset.ai and requires Node.js 22.12.0 or newer and Bun 1.3.6. Workspaces are organized under apps/, packages/, and tooling/.

The web application is the Next.js frontend and API layer, built with React and Tailwind CSS v4. The AI SDK referenced in the README is Vercel's AI SDK, which handles LLM provider abstraction. Prisma manages the database schema and migrations against a PostgreSQL instance. Trigger.dev handles background job processing, particularly for asynchronous document ingestion. Supabase provides the authentication backend.

The monorepo includes apps/, packages/, tooling/, and docs/ directories at the top level, along with configuration for .claude/, .cursor/, .zed/, and .opencode/ reflecting integration with several AI development tools. This layout suggests the codebase is actively developed with AI coding assistants, which is consistent with the last push on 2026-09-27.

The model-agnostic claim in the README refers to the fact that Agentset does not hard-code a specific LLM or embeddings provider. The .env.example uses Azure for managed models by default, but the slot is configurable. The vector store is similarly configurable between Pinecone and Turbopuffer depending on which API key is provided.

Getting Agentset Running Locally

The README documents a four-step local development setup.

First, copy the environment file and fill in the required values:

bash
cp .env.example .env

Second, install dependencies:

bash
bun install

Third, run database migrations:

bash
bun db:deploy

Fourth, start the web application:

bash
bun dev:web

The README also lists bun db:studio for opening Prisma Studio, which provides a GUI for inspecting the database schema and records.

The .env.example sets DATABASE_URL to a local PostgreSQL connection string. This means a PostgreSQL instance must be running before the migration step. The README links to a full self-host prerequisites guide at docs.agentset.ai/open-source/prerequisites for teams who need to set up all dependencies from scratch.

The cloud alternative requires no local setup: signing up at app.agentset.ai provides access to the hosted version with a free tier that includes 1,000 pages and 10,000 retrievals, requiring no credit card.

File Format Support, Partitions, and the Ingestion Pipeline

The project description references 22 or more file formats and partitions. The .env.example documents a PARTITION_API_URL and PARTITION_API_KEY for the partitioning service that handles document ingestion. This is a separate API, external to the main application, responsible for parsing and extracting text from uploaded files before they are chunked and indexed.

The README describes the ingestion pipeline as turnkey: ingestion, chunking, embeddings, and retrieval are handled by the platform. S3-compatible object storage holds the raw uploaded files, the partitioning API extracts text, the embedding service vectorizes the chunks, and the vector store (Pinecone or Turbopuffer) indexes them for retrieval.

The MCP server referenced in the project description means Agentset exposes its retrieval capabilities as a Model Context Protocol server, allowing MCP-compatible AI clients to query indexed document collections directly. The README does not give a separate section on MCP configuration, so teams that specifically need this capability should check the documentation at docs.agentset.ai.

Multi-tenancy is listed as a built-in feature, meaning a single Agentset deployment can serve multiple customers or teams with isolated document collections. This is handled through the partition concept, which appears in both the environment variable PARTITION_API_URL and the project description's feature list.

Self-Hosting: The Full Dependency Surface

The .env.example documents every external service a self-hosted Agentset requires. The required infrastructure spans several categories.

Database: PostgreSQL at the connection string in DATABASE_URL.

Authentication: BETTER_AUTH_SECRET and BETTER_AUTH_URL for the Better Auth library, plus optional GitHub and Google OAuth credentials.

Vector store: Either DEFAULT_PINECONE_API_KEY and DEFAULT_PINECONE_HOST for Pinecone, or DEFAULT_TURBOPUFFER_API_KEY for Turbopuffer.

Embeddings and re-ranking: DEFAULT_AZURE_RESOURCE_NAME and DEFAULT_AZURE_API_KEY for Azure-hosted models, plus DEFAULT_COHERE_API_KEY and DEFAULT_ZEROENTROPY_API_KEY for re-ranking.

Document ingestion: PARTITION_API_URL and PARTITION_API_KEY for the external partitioning service.

File storage: S3_ACCESS_KEY, S3_SECRET_KEY, S3_ENDPOINT, and S3_BUCKET for any S3-compatible object store.

Cache and queuing: REDIS_URL and REDIS_TOKEN for Upstash Redis.

Payments: STRIPE_WEBHOOK_SECRET and STRIPE_API_KEY for Stripe billing integration.

Analytics: TINYBIRD_API_KEY and TINYBIRD_API_URL for Tinybird analytics.

This is a significant operational footprint. Teams self-hosting for data residency control will need to provision or substitute each of these services before the application starts. Missing any required variable will prevent the application from running correctly.

Where Agentset Is the Wrong Tool

Teams who need a lightweight RAG prototype in an afternoon will find the self-host dependency list prohibitive. The six-plus external services in .env.example are appropriate for a production deployment but represent substantial setup overhead for evaluation purposes. The cloud tier at app.agentset.ai exists specifically for this case.

Teams that need to use a vector store other than Pinecone or Turbopuffer will need to extend the codebase, since those are the two explicitly configured options in .env.example.

The repository has no GitHub releases, and the package.json shows version 0.0.0, suggesting the project has not yet adopted a formal versioning scheme. This means there are no stability guarantees for the API or the database schema between updates, and teams self-hosting will need to check for breaking schema migrations before updating.

Ragie is a managed RAG-as-a-service platform and Morphik AI is another managed document intelligence service that both address similar use cases. Both are hosted services with their own data processing pipelines, rather than self-hostable open-source platforms. The distinction is relevant for teams with compliance requirements that prohibit sending documents to third-party APIs: Agentset's self-host path keeps documents within the operator's own infrastructure.

Editorial conclusion

Teams building production RAG applications who want an open-source codebase they can inspect and extend are the right audience. The cloud tier at app.agentset.ai provides a free starting point without configuring any external services, and the self-host path documents the full .env.example for teams that need data residency control. Before self-hosting, verify that PostgreSQL, a vector store (Pinecone or Turbopuffer), Upstash Redis, and S3-compatible storage are all available in your infrastructure, because the application will not start without them.

Frequently asked questions

How do I start Agentset for local development?

The README documents four commands: copy .env.example to .env and fill in required values, run bun install, run bun db:deploy for database migrations, then run bun dev:web to start the web application. A running PostgreSQL instance and the other services in .env.example are required before the app will start.

Does Agentset require a specific vector store?

The .env.example documents Pinecone and Turbopuffer as the two vector store options. The environment variables for each are mutually exclusive: Agentset uses whichever one has credentials populated in the .env file. The README does not document adding a third vector store without modifying the codebase.

Is there a free way to try Agentset without self-hosting?

Yes. The README documents a cloud tier at app.agentset.ai that provides 1,000 pages and 10,000 retrievals on a free tier with no credit card required. This is the fastest path to evaluating the platform before committing to a self-hosted setup.

Official sources

  1. agentset-ai/agentset on GitHub
  2. Issues
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/agentset-ai-agentset.svg)](https://hysenlabs.com/projects/agentset-ai-agentset)