Model or dataset
microsoft/kernel-memory avatar
microsoft/kernel-memory

Kernel Memory: Microsoft's RAG Reference Implementation, Archived and Unsupported

Research project. A Memory solution for users, teams, and applications.

2,235 stars415 forksC#MIT

At a glance

What is it?
Kernel Memory is a multi-modal indexing and RAG service from Microsoft, shipped as a .NET library, a web service and a Docker image. The repository itself calls it an archived research project, so the question is not whether it works but who should still build on it.
Who is it for?
Adopt Kernel Memory if you are prototyping RAG on .NET, want a working reference for pipelines, tags and citations, and accept that no support is provided. Do not adopt it if you need a maintained dependency with a support contract, or if your stack is Python or TypeScript and you would only be calling the HTTP service.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 113 days ago.
What is it written in?
Mainly C#, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Kernel Memory actually is, and the warning at the top of the README

Kernel Memory is a memory service for AI and LLM applications. The repository describes it as a multi-modal AI service specialized in indexing datasets through custom continuous data hybrid pipelines, with support for Retrieval Augmented Generation, synthetic memory, prompt engineering and custom semantic memory processing. It is not a vector database and not an agent framework. It sits between your documents and your model, handling the work of turning files into retrievable chunks and turning questions into grounded answers with citations.

The README opens with a caution block that is unusual for a Microsoft repository: this is an archived research project, the code serves as a learning resource rather than production software, no support is provided, and use is at your own risk. A later paragraph repeats that the code is a demonstration and not an officially supported Microsoft offering. That framing should drive the adoption decision more than any feature list. The last push to the repository was on 2026-06-08, so it is not abandoned in the sense of a dead fork, but the project's own description tells you what kind of dependency it is.

The audience is narrower than the feature list suggests. It is for .NET engineers who want to see how a RAG pipeline is assembled end to end, and for teams evaluating architectures before they commit to a vendor. The examples directory alone is a substantial reference: custom pipelines, custom embedding generators, custom LLM connectors, Anthropic and LlamaSharp integrations, an NL2SQL example, an ASP.NET MVC integration, and a text extraction example. Reading those is a legitimate reason to clone the repository even if you never deploy it.

The ingestion pipeline: extract, partition, embed, index

The README spells out four default stages for document ingestion. First, extract text by automatically recognizing the file format. Second, partition the text into small chunks ready for search and RAG prompts. Third, extract embeddings using any LLM embedding generator. Fourth, save embeddings into a vector index such as Azure AI Search or Qdrant. That is the whole data flow, and it is exposed rather than hidden, which is why the custom pipeline examples exist.

Two design choices stand out. Documents carry tags, and tags are the mechanism for both organization and access control. The README's own example applies user, collection and fiscalYear tags to a business plan, and the retrieval example passes a filter to restrict a question to one user's documents. The README frames this as safeguarding private information, but tags are metadata, not authentication. If the caller can choose the filter, the caller can choose to omit it. Isolation has to be enforced by whatever wraps the service, not by Kernel Memory itself.

The second choice is that the same abstractions work across deployment shapes. A KernelMemoryBuilder with OpenAI defaults builds a MemoryServerless instance for embedded use; a MemoryWebClient pointed at a URL talks to the service. The C# import call looks the same in both cases. That symmetry is the most reusable idea in the repository, and it is the part most worth copying into your own design even if you do not take the library.

Installing Kernel Memory and asking your first question

The README points at three distribution paths: a .NET library for embedded applications, a web service, and a Docker image published as kernelmemory/service on Docker Hub. There is no NuGet install command printed in the README, but the C# examples reference the package Microsoft.KernelMemory.WebClient, and the repository contains a nuget.config and a Directory.Packages.props, so the packages are published under the Microsoft.KernelMemory namespace.

For the embedded path, the README gives this builder, which wires OpenAI as the default text and embedding provider using an environment variable for the key:

csharp
var memory = new KernelMemoryBuilder()
    .WithOpenAIDefaults(Environment.GetEnvironmentVariable("OPENAI_API_KEY"))
    .Build<MemoryServerless>();

Once built, importing a file and applying tags uses the same call shape as the web client. The Document object takes an ID and a chain of AddTag calls:

csharp
await memory.ImportDocumentAsync("business-plan.docx",
    new Document("doc01")
        .AddTag("collection", "business")
        .AddTag("collection", "plans")
        .AddTag("fiscalYear", "2025"));

Questioning is a single call, optionally filtered by tag. The README's example filters on the user tag to scope an answer to one person's documents:

csharp
var answer = await memory.AskAsync("what's the project timeline?",
    filter: MemoryFilters.ByTag("user", "[email protected]"));

If you would rather not write C#, the service exposes an HTTP endpoint. The README's Python example posts to http://127.0.0.1:9001/upload with a files dictionary and a data dictionary containing documentId and a list of tags in the form collection:business. The C# web client is constructed against the same host and port: new MemoryWebClient("http://127.0.0.1:9001"). Port 9001 is what both examples use.

For the container path, the README documents a .NET Aspire setup that adds the kernelmemory/service container and configures it entirely through environment variables, including KernelMemory__TextGeneratorType set to OpenAI, KernelMemory__DataIngestion__EmbeddingGeneratorTypes__0 set to OpenAI, KernelMemory__Retrieval__EmbeddingGeneratorType set to OpenAI, and KernelMemory__Services__OpenAI__APIKey for the key. That double-underscore convention is how the .NET configuration system maps nested settings, and it is the same set of options whether you set them in Aspire, in a compose file or in the service's own configuration.

Answers come with citations and a token bill

The retrieval side is where Kernel Memory tries to be more than a search wrapper. The README states that answers include citations and all the information needed to verify their accuracy, pointing to which documents ground the response. For anything user-facing, that is the difference between a demo and something a reviewer will accept, because a citation lets a reader check the claim rather than trust it.

The second retrieval feature is a token usage report. When answers are generated with LLMs, the result carries per-service accounting, and the README shows a loop over the reports printing ServiceType, ModelName, ModelType, ServiceTokensIn and ServiceTokensOut. The sample output it prints is Azure OpenAI gpt-4o with 24356 input tokens and 103 output tokens. That ratio is the point: retrieval is cheap relative to the prompt you assemble from it, and a pipeline that retrieves too many chunks will spend most of its budget on input tokens. If you are sizing a deployment, that report is the instrument you need, and it is built into the answer rather than bolted on.

Neither feature is a guarantee of quality. Citations tell you which chunks were used; they do not tell you whether the chunking was sensible, and the README does not document an evaluation harness for measuring whether answers are correct. You would be building that yourself.

Where Kernel Memory is the wrong tool

The most important limitation is stated by the project, not inferred: no support is provided, and the code is not production software. If your organization requires a vendor to answer a support ticket, this repository cannot satisfy that requirement, and no amount of feature coverage changes it. The MIT licence means you can fork and maintain it yourself, but that is a commitment to owning a RAG pipeline, not a way to avoid one.

The second limitation is the tag-based filtering model described above. The README presents tags as a way to safeguard private information, and the retrieval example shows filtering by user. But filtering happens at query time on metadata the caller supplies, so multi-tenant isolation depends entirely on the layer in front of the service. A team that reads the README quickly could ship a service where any caller can query any document by omitting the filter.

The third is configuration surface. The Aspire example alone sets four environment variables just to select OpenAI for text generation, embeddings and retrieval. Every additional provider, store or pipeline stage adds more of the same. That is normal for a service that integrates with many backends, but it means the operational burden scales with the number of moving parts you enable, and the README does not document a rollback or migration path when you change an embedding model. Changing embeddings invalidates the index, and nothing in the project's documentation describes how to handle that transition.

Finally, the name itself is a problem. Kernel memory is a term of art in operating systems for memory owned by the kernel, and the related searches around it are dominated by that meaning: kernel memory dumps, kernel memory leaks, kernel memory allocation in OS, kernel memory vs user memory. If you are searching for help, expect to filter past a large amount of operating system material. Adding Microsoft or the GitHub repository name to the query is the practical workaround.

Kernel Memory against Semantic Kernel, and against a plain vector store

The comparison people actually search for is Kernel Memory versus Semantic Kernel, and the two are not competitors. Semantic Kernel is the orchestration framework for building agents and calling models; Kernel Memory is the retrieval and indexing layer. The README says Kernel Memory is designed to integrate as a plugin with Semantic Kernel, Microsoft Copilot and ChatGPT, and the repository includes an example directory for the Semantic Kernel plugin. If you already use Semantic Kernel, Kernel Memory is a component you can attach, not a replacement.

The more interesting comparison is against assembling the pipeline yourself on top of a vector store. A vector database gives you storage and similarity search; it does not give you format-aware text extraction, partitioning, a document model with tags, or citation-bearing answers. Kernel Memory provides those stages and lets you swap the underlying store, with Azure AI Search and Qdrant named in the README. The trade is control. If your documents are one format and your chunking strategy is fixed, a direct integration with a vector store is fewer moving parts. If you have mixed formats and want a starting shape for the pipeline, the reference implementation saves you design time, and its examples show how to replace individual stages when the defaults do not fit.

Editorial conclusion

Adopt Kernel Memory if you are prototyping RAG on .NET, want a working reference for pipelines, tags and citations, and accept that no support is provided. Do not adopt it if you need a maintained dependency with a support contract, or if your stack is Python or TypeScript and you would only be calling the HTTP service. Before committing, verify that the pipeline stages you need (text extraction, partitioning, embedding, vector store) are covered by the examples in the repository, and read the archived notice in the README, which states the code is a learning resource and not production software.

Frequently asked questions

What is Kernel Memory from Microsoft?

It is a multi-modal AI service for indexing datasets through custom continuous data hybrid pipelines, with support for Retrieval Augmented Generation and semantic memory. The README describes it as a reference implementation and demonstration, not an officially supported Microsoft offering.

How do I install Kernel Memory?

The README lists three paths: a .NET library for embedded applications, a web service, and a Docker image published as kernelmemory/service. The C# examples reference the package Microsoft.KernelMemory.WebClient, and the service examples run on port 9001.

What is the difference between Kernel Memory and Semantic Kernel?

Kernel Memory handles indexing and retrieval with citations, while Semantic Kernel is the framework it plugs into. The README states Kernel Memory is designed for integration as a plugin with Semantic Kernel, Microsoft Copilot and ChatGPT.

What does Kernel Memory store?

It stores the text extracted from your documents, partitioned into chunks, together with the embeddings generated for those chunks in a vector index such as Azure AI Search or Qdrant. Tags attached to each document are stored alongside and used to filter queries.

Where is Kernel Memory stored?

The README names Azure AI Search and Qdrant as vector index options for the embeddings, and the service itself is distributed as a Docker image, kernelmemory/service, as well as a .NET library and a web service. The README does not document a single fixed storage location.

Official sources

  1. Issues
  2. License: MIT
  3. microsoft/kernel-memory on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/microsoft-kernel-memory.svg)](https://hysenlabs.com/projects/microsoft-kernel-memory)