nextjs-openai-doc-search: a Supabase RAG template for your own docs
Template for building your own custom ChatGPT style doc search powered by Next.js, OpenAI, and Supabase.
At a glance
- What is it?
- Supabase's starter turns .mdx pages into a pgvector index and answers questions through an edge function. It is a template for teams already on Supabase, not a general purpose search product.
- Who is it for?
- Adopt it if your documentation already lives in .mdx files inside a Next.js repository and you are willing to run Supabase, because the embedding script, the migration and the edge function assume exactly that layout. Do not adopt it if you need a crawler for external sites, a managed search backend, or a retrieval layer you can swap without touching application code.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 141 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem it solves: docs that answer questions instead of listing links
Keyword search returns a list of pages and leaves the reading to you. This starter returns a written answer assembled from your own documentation. It targets a narrow audience: teams whose docs are already .mdx files sitting in a Next.js repository, and who want a ChatGPT style box in front of them. The README frames the whole thing as a starter, and that word is accurate. There is no admin interface, no crawler, no ingestion pipeline for external sources. If your documentation lives in Confluence, a Notion workspace or a separate docs site, this template gives you nothing until you convert that content into .mdx files under pages. The payoff for the target audience is real: the retrieval index, the database schema and the streaming chat UI all ship in the repository, so the work left is content preparation and an OpenAI key.
Build time versus runtime: where the four steps actually happen
The README splits the system into four steps and marks two of them build time and two runtime. Build time covers preprocessing and embedding. The generate-embeddings script in lib chunks each .mdx page into sections, asks OpenAI for an embedding per section, and writes the vector into Postgres through pgvector. The README states the vector has 1536 dimensions. Alongside the vectors, the script stores a checksum per .mdx file in a separate table, so a rebuild only regenerates embeddings for files whose contents changed. That checksum table is the difference between a cheap redeploy and paying OpenAI to re-embed your entire documentation on every push.
Runtime starts when a user submits a question. The SearchDialog component posts to the vector-search edge function, which embeds the query, runs a similarity search against pgvector, and then sends a completion request to OpenAI with the query and the retrieved content combined into one prompt. The response streams back to the client as text/event-stream. The README labels both runtime steps as critical in its sequence diagram, which is a fair description: if the embedding call fails, retrieval has no vector to compare against, and the answer degrades to whatever the model already knows.
Installing it locally and running your first query
Local development needs Docker running, because Supabase starts as containers. Copy the environment file first, then fill in the three secrets. The example file lists the local Supabase URL as http://localhost:54321 and leaves the anon key, the service role key and the OpenAI key blank.
cp .env.example .envThe README notes that the two Supabase keys come from a running instance, so start Supabase before you look for them. The migration in supabase/migrations, which enables the pgvector extension, is applied automatically when the stack comes up.
supabase start
supabase statusThe status command prints the values you paste into NEXT_PUBLIC_SUPABASE_ANON_KEY and SUPABASE_SERVICE_ROLE_KEY. With the keys in place, install dependencies and start Next.js.
pnpm install
pnpm devThe app serves on localhost:3000. To index your own content, rename or convert your markdown to .mdx under pages, then run the embedding script and restart the dev server so the rendered page picks up the new index.
pnpm run embeddings
pnpm devThe package.json exposes a second script, embeddings:refresh, which passes a --refresh flag to the same tsx entry point. The README does not document what that flag changes, so treat it as the escape hatch for forcing a rebuild rather than a documented workflow.
The build step is a hard dependency, not a background job
The build script is defined as pnpm run embeddings && next build. Embedding generation is therefore inside the deployment path, not beside it. Every production build calls OpenAI once per changed section, and if that call fails or the API key is missing, the build fails with it. For a documentation site that deploys several times a day, that means your release cadence is coupled to an external API's availability and to your remaining quota. The checksum table limits the blast radius to files that actually changed, but a large edit still produces a burst of embedding requests in the middle of a build. Teams used to decoupling content ingestion from application deploys will find this arrangement uncomfortable, and the README does not describe a way to run the embedding step out of band.
What it does not cover: auth, rate limits and non-mdx sources
The repository is a starter, and the gaps are the ones you would expect. There is no authentication layer around the API route, so anyone who can reach the deployment can spend your OpenAI credits. There is no rate limiting, no caching of repeated questions, and no evaluation harness for checking whether retrieved sections are actually relevant. The README does not document rollback, index versioning or a way to query the vector store without going through OpenAI. Content support stops at .mdx: the preprocessing pipeline parses markdown with mdast and the MDX extensions, so HTML pages, PDFs and OpenAPI specs are out of scope until you write your own ingestion code. None of this is a defect in a template whose stated job is to demonstrate the pattern, but it does mean the distance between this repository and a production documentation assistant is measured in application code you write yourself.
How it differs from a managed search service
A hosted search product such as Algolia or Typesense indexes documents through its own API and returns ranked hits, and you decide how to present them. This template takes the opposite position: retrieval and generation are one pipeline, and the output is prose. That buys you answer synthesis and costs you the ability to tune ranking without touching the prompt. The other real alternative is assembling the same stack by hand, using pgvector directly and a general purpose RAG framework for chunking and retrieval. The template's advantage there is that the chunking logic, the checksum table, the migration and the streaming edge function are already wired together and consistent. Its disadvantage is the same thing: the pieces are coupled, so replacing the chunker or the embedding model means editing the script rather than swapping a plugin. If you want a search backend you can point any client at, this is the wrong shape. If you want a working reference for Supabase plus OpenAI retrieval, the coupling is the point.
Licence, maintenance and the cost of upgrading
The repository is Apache-2.0, which permits commercial use and modification, and requires that you keep the licence and notice files when you redistribute. It includes no warranty, so the correctness of answers generated from your documentation is your problem, not the project's. Nothing here is legal advice; check how Apache-2.0 interacts with your own distribution model.
On maintenance, the last push to the default branch was on 2026-05-12, and the repository is not archived. The dependency list pins Next.js 13.2.4, React 18.2.0 and the openai package at ^3.3.0, with openai-edge alongside it for the streaming edge function. Those are older major lines. Upgrading means reconciling the AI SDK version used by the streaming route, the OpenAI client version in the embedding script, and the edge runtime's constraints, and none of that is covered by the README. Budget for a dependency pass before you build on top of it, and expect the migration SQL under supabase/migrations to be the stable part of the repository.
Editorial conclusion
Adopt it if your documentation already lives in .mdx files inside a Next.js repository and you are willing to run Supabase, because the embedding script, the migration and the edge function assume exactly that layout. Do not adopt it if you need a crawler for external sites, a managed search backend, or a retrieval layer you can swap without touching application code. Before committing, verify three things in your own checkout: that every .mdx file parses through the mdast pipeline, that your OpenAI account can pay for one embedding call per changed section on each build, and that you can regenerate the index locally with pnpm run embeddings:refresh after editing a page.
Frequently asked questions
What is nextjs-openai-doc-search and who is it for?
It is a template for building a ChatGPT style search over your own documentation, using Next.js, OpenAI and Supabase. It suits teams whose docs are already .mdx files in a Next.js repository and who are willing to run Supabase with pgvector.
How do I install nextjs-openai-doc-search locally?
Copy .env.example to .env, run supabase start with Docker running, take the keys from supabase status, then install dependencies and run pnpm dev. The app serves on localhost:3000, and the README lists the local Supabase URL as http://localhost:54321.
How does nextjs-openai-doc-search handle my .mdx documentation?
The build runs the generate-embeddings script, which chunks each .mdx page into sections, creates an embedding per section through the OpenAI API, and stores the 1536 dimension vector in Postgres with pgvector. A checksum per file is stored so embeddings are only regenerated when a file changes.
What does nextjs-openai-doc-search do at runtime when a user asks a question?
The SearchDialog component sends the query to the vector-search edge function, which embeds the query, runs a vector similarity search in pgvector, and injects the retrieved content into an OpenAI completion prompt. The answer streams back to the client as text/event-stream.
Is nextjs-openai-doc-search free to use?
The code is licensed under Apache-2.0, so there is no licence fee, but running it requires a Supabase instance and an OpenAI API key, and the build script calls the OpenAI API for embeddings on changed files.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/supabase-community-nextjs-openai-doc-search)