Open-source project
SaiAkhil066/CORTEX-AI-SUPER-RAG avatar
SaiAkhil066/CORTEX-AI-SUPER-RAG

CORTEX RAG: nine retrieval techniques stacked into one local document assistant

CORTEX RAG is an enterprise retrieval and knowledge assistant that helps teams find accurate answers from company data with citations, permission-aware retrieval, and fast deployment

1,977 stars290 forksPythonMIT

At a glance

What is it?
A Streamlit app that keeps a PDF, the model and the index on your own machine, and layers contextual retrieval, query fusion, a knowledge graph, reranking and corrective grading into one pipeline. The install step pins half its dependencies to git branches, which is the part to read before you run it.
Who is it for?
CORTEX RAG does something most RAG demos do not: it names each retrieval technique it uses and draws the flow diagram that shows where they sit, so you can argue with the architecture rather than guess at it. That makes it a good teaching artifact and a reasonable starting skeleton.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 11 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Cloning, installing and pulling two Ollama models

The quick start is four steps and two prerequisites: Ollama installed and running, and Python 3.10 or higher available. The clone is the usual pair of commands.

bash
git clone https://github.com/SaiAkhil066/CORTEX-AI-SUPER-RAG.git
cd CORTEX-AI-SUPER-RAG

Dependencies come from a requirements file, and then two models are pulled. The embedding model is described as required; the language model is described as swappable.

bash
pip install -r requirements.txt
bash
ollama pull llama3.1:8b          # LLM  (swap for any model you prefer)
ollama pull nomic-embed-text     # Embeddings  (required)

The app is a Streamlit server, and the README is insistent about how you launch it:

bash
python -m streamlit run app.py

Port 8501, and the README explains why the module form matters, saying `python -m streamlit run` rather than bare `streamlit run` ensures the correct Python environment is picked up. That is a real failure mode on a machine with several interpreters, and it is the kind of detail most READMEs skip.

There is one platform-specific escape hatch. On Windows, a `c10.dll` error on first run is addressed by uninstalling torch and reinstalling a pinned CPU build, which is a symptom of what the next section is about.

Nine libraries resolved from GitHub default branches

The requirements file is the most consequential file in this repository, and not for the reason you would expect. Twenty packages are listed. Eighteen of them are ordinary name pins you would see anywhere. Nine are not pinned at all: `langchain`, `langchain-community`, `langchain-core`, `langchain-classic` and `langchain-ollama` appear with no version constraint.

The other problem is that the core numeric and machine learning stack is installed from source control rather than from an index:

code
torch @ git+https://github.com/pytorch/pytorch.git
numpy @ git+https://github.com/numpy/numpy.git
pandas @ git+https://github.com/pandas-dev/pandas.git
faiss-cpu @ git+https://github.com/facebookresearch/faiss.git
streamlit @ git+https://github.com/streamlit/streamlit.git

The same pattern covers `requests`, `networkx`, `sentence-transformers`, `rank-bm25`, `pydantic`, `PyMuPDF`, `docx2txt`, `tqdm`, `pypdf` and `python-dotenv`, each pointing at its GitHub repository with no ref, tag or commit.

Read plainly, `torch @ git+https://github.com/pytorch/pytorch.git` installs whatever is on PyTorch's default branch at the moment you run pip, which is a nightly or a release candidate rather than a version anyone certified against this code. Building torch, FAISS and sentence-transformers from source also means a working C and C++ toolchain, and it turns a two minute install into something closer to an afternoon on most machines.

There is a second-order effect worth naming. Because the pinned versions come from moving branches rather than fixed commits, two people running `pip install -r requirements.txt` on different days can end up with genuinely different stacks and different failures. That is the practical reason to fork this and pin the refs yourself before treating it as a base for anything.

The one saving grace is that the Docker path sidesteps the host toolchain entirely, at the cost of the Dockerfile still doing a bare `pip install --no-cache-dir -r requirements.txt`, so it pulls the same moving branches inside the container.

Three indexes, then fusion, reranking and grading

The README's flow diagram is the best thing in the documentation, and it earns the word architecture more than the prose around it does. The pipeline has a clear spine.

Documents are chunked. If contextual retrieval is on, an LLM enriches every chunk with its surrounding context before indexing, so a vector carries the local story of the document rather than an orphan fragment. Three indexes are then built in parallel: BM25 for lexical matching, FAISS for dense vectors, and a NetworkX graph over document entities. When a query arrives it first hits a semantic cache, then fans out into RAG-Fusion, which generates several query variants, retrieves independently for each, and merges the ranked lists with Reciprocal Rank Fusion. A graph entity boost is applied on top. A cross-encoder reranker then reorders everything by true query-passage relevance. Finally, corrective RAG has the model grade each surviving chunk for relevance and silently drop the noise before generation begins.

Two details in there are the ones most implementations skip. Grading before generation, rather than after, means the model never sees a chunk it was going to discard, which changes both cost and hallucination surface. And reranking by a cross-encoder rather than by embedding similarity catches the case where a passage shares vocabulary with the question but does not answer it.

The named reranker is `ms-marco-MiniLM`, and the compose file confirms the exact identifier as `cross-encoder/ms-marco-MiniLM-L-6-v2`, a small model, which is consistent with running the whole thing on a laptop.

The response layer streams the model's `<think>` chain-of-thought live in a panel while the answer and source cards render at the end, and full multi-turn history flows into every generation call so follow-up questions work.

A semantic cache that answers on cosine similarity alone

The ninth technique deserves its own look because it is the one that changes behaviour in a way users notice. The cache is a cosine-similarity check at a threshold of 0.92 on the query embedding. Above that threshold, the pipeline skips retrieval and skips generation entirely and returns the earlier answer.

The upside is large in exactly the situation this app is used in. Someone asking the same question about the same document three times gets the first answer back three times, and the second and third cost a fraction of the first. For a 8B model on local hardware, that is the difference between a responsive tool and a spinner.

The risk is that the threshold is doing a lot of work on its own. Cosine similarity between two query embeddings at 0.92 does not mean the questions want the same answer. Two queries that differ by one word, a date, or a constraint can sit above that line, and the cached answer would then be confidently wrong for the question actually asked. Because the cache is described as keyed on the query embedding rather than on the query together with the uploaded document, a second document indexed after the first raises a further question about whether an earlier answer can be served against a later corpus.

The README does not describe invalidation, per-document namespacing or an eviction policy, and it does not say what happens when the index changes underneath a cached answer. Treat 0.92 as a tunable starting point rather than a settled number, and find where it is defined before you rely on it.

Permission-aware retrieval is promised in the description and absent from the code

The repository description on GitHub calls this an enterprise retrieval and knowledge assistant that helps teams find accurate answers from company data with citations, permission-aware retrieval, and fast deployment. The README's body pitch is a local, offline tool: upload a PDF, ask a question, no API key, no cloud upload, no subscription, with a badge reading zero-cloud.

Neither description matches what the pipeline does. Permission-aware retrieval means the system knows which documents a given user is allowed to see and filters retrieval by that. Nothing in the README, the flow diagram, the requirements file or the docker-compose configuration describes an access control layer, an identity, a session or a per-user filter. The pipeline indexes everything uploaded and retrieves across all of it.

This is the one place in this project where the gap between the pitch and the repository is large enough to matter, so it is worth being exact about why. A single-user local tool over a document the user already possesses does not need an ACL layer, and its absence is defensible. What the description implies, several users querying company data and each seeing only their own slice, is a different system entirely, and the README does not describe one.

There is a related wrinkle: the README devotes a prominent block to offering custom production RAG builds through a separate hosted site, and the stars badge in the header points at a different repository, `DeepSeek-RAG-Chatbot`, rather than this one. Both details suggest a portfolio project that has been reshaped, and neither is a problem for reading the code, but both are reasons to treat the enterprise framing as a business advertisement rather than a feature list.

What the container config does and does not solve

The Dockerfile is short and has not been cleaned up. It builds from `python:3.11-buster` with a comment telling you to change buster to slim, sets a maintainer label to `[email protected]`, copies the requirements file, installs with `--no-cache-dir`, exposes 8501 and runs `streamlit run app.py`. The placeholder label is a small thing that dates the file to a template.

The compose file is where the interesting configuration lives. It maps 8501 to the host, then sets four environment variables that together describe the entire model stack:

yaml
OLLAMA_API_URL=http://host.docker.internal:11434
MODEL=llama3.1:8b
EMBEDDINGS_MODEL=nomic-embed-text:latest
CROSS_ENCODER_MODEL=cross-encoder/ms-marco-MiniLM-L-6-v2

That first line is the detail that makes the container usable rather than merely present. Pointing at `host.docker.internal` on port 11434 is how the container reaches an Ollama server running on your laptop, which keeps the models off the image and off disk inside Docker. A live-development volume mount and a custom command to bind the server to `0.0.0.0` are both present but commented out.

So the container does solve the toolchain problem from the requirements file, because a Docker image can carry a compiler toolchain that a laptop should not have to. It does not solve the moving-branch problem, since the same unpinned requirements are installed inside it, and the compose file still declares `version: "3.8"`, a field modern Compose ignores.

Editorial conclusion

CORTEX RAG does something most RAG demos do not: it names each retrieval technique it uses and draws the flow diagram that shows where they sit, so you can argue with the architecture rather than guess at it. That makes it a good teaching artifact and a reasonable starting skeleton. Two things bound the enthusiasm. The requirements file pulls torch, numpy, pandas, streamlit, faiss-cpu, pydantic, PyMuPDF and eight more straight from GitHub default branches, so there is no reproducible environment and a breaking upstream commit reaches you on the next install. And the project description promises permission-aware retrieval for enterprise data while nothing in the code or docs describes an access control layer, which is the gap between the marketing page and the repository. For a local read-and-cite tool over documents you already trust, run it. For anything multi-user, the pinning and the missing ACL layer are what to solve first.

Frequently asked questions

What is CORTEX RAG and how does it answer questions from a PDF?

It is a local Streamlit app that runs retrieval over documents you upload, with no cloud call and no API key. The pipeline builds BM25, FAISS and a NetworkX graph index, then on each query tries a cosine cache at 0.92, expands the query into variants and merges them with Reciprocal Rank Fusion, reranks with an ms-marco-MiniLM cross-encoder, and has the model grade chunks before generating a cited answer.

How do I install and run CORTEX RAG locally?

Install Ollama and Python 3.10 or higher, clone the repository, run `pip install -r requirements.txt`, then `ollama pull llama3.1:8b` and `ollama pull nomic-embed-text` before starting the app with `python -m streamlit run app.py` and opening port 8501. The README asks you to use the `python -m` form so the right interpreter is picked up.

Does CORTEX RAG work without internet access?

Yes, once the models are pulled. Inference runs through Ollama on your machine and the embedder is `nomic-embed-text` served locally, and the README describes no external API call in the query path. The only network need is the initial model download and the dependency install.

Why does installing CORTEX RAG take so long or fail on Windows?

The requirements file installs torch, numpy, pandas, FAISS, Streamlit, PyMuPDF, sentence-transformers and several more straight from their GitHub default branches with no ref, so pip builds them from source and can pick up an untested upstream commit. The README also documents a Windows-only fix, uninstalling torch and reinstalling a pinned CPU build when a `c10.dll` error appears.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. SaiAkhil066/CORTEX-AI-SUPER-RAG on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/saiakhil066-cortex-ai-super-rag.svg)](https://hysenlabs.com/projects/saiakhil066-cortex-ai-super-rag)