EnterpriseRAG-Bench: a 500,000-document synthetic company for testing RAG
Dataset and benchmark for RAG on company internal documents.
At a glance
- What is it?
- Onyx's EnterpriseRAG-Bench ships a synthetic company called Redwood Inference, nine internal source types and 500 graded questions. It is a retrieval benchmark, not a full answer-quality benchmark, and that distinction decides who should use it.
- Who is it for?
- Adopt EnterpriseRAG-Bench if you are choosing or tuning a retriever over heterogeneous internal sources and need ground-truth document IDs, and treat the 40 High Level plus 20 Info Not Found questions as your abstention check. Do not adopt it as your only evaluation if your product lives or dies on answer phrasing, or if you need a corpus in a non-English language: the README describes an English-language synthetic company and does not document multilingual generation.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 26 days ago.
- What is it written in?
- GitHub does not report a main language for this repository.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap EnterpriseRAG-Bench fills: internal corpora, not public web text
Public RAG benchmarks lean on text anyone can download: web search results, developer forums, open Q&A sites. The README states the project's premise plainly: existing RAG and IR datasets largely focus on publicly accessible document sets, and there had not been a publicly accessible dataset focused entirely on company internal data. EnterpriseRAG-Bench targets that gap with a synthetic company, Redwood Inference, which sells AI model inference as a service. The corpus is slightly over 500,000 documents spread across nine source types, with Slack at roughly 275,000 documents and Gmail at 120,000 doing most of the volume work, down to Confluence at 5,000. The audience is narrow on purpose: teams building or procuring a RAG system for employee-facing search, and researchers who need a corpus where the hard part is the corpus itself. Retrieval over a wiki is easy. Retrieval over 275,000 Slack messages, near-duplicates, misfiled documents and project codenames is the actual job.
How the synthetic company is generated: scaffolding first, volume second
The generation process is documented in methodology.md and summarised in the README's five principles. It starts with human-in-the-loop scaffolding: a company overview, initiatives, an employee directory and a source structure. A core set of high-fidelity documents is then produced with awareness of that context and of each other, which is what gives the corpus cross-document coherence. Only after that does high-volume generation run, using topic-based scaffolding to keep diversity while controlling cost.
Noise is a deliberate design decision rather than an artefact. The README describes introducing noise through random and LLM-driven document shuffling, miscellaneous files, and near-duplicates carrying updated or conflicting facts. Documents also use project codenames, product acronyms and organisational jargon that are meaningless outside the company but required for correct retrieval. That combination is what makes the benchmark uncomfortable in a useful way: a retriever that keys on lexical overlap will fail the Semantic category, and one that trusts the first plausible hit will fail the Conflicting Info and Constrained categories.
The 500 questions, category by category
The question set lives in questions.jsonl and totals 500 items across ten categories. Basic (175) and Semantic (125) are the bulk: single ground-truth documents, with Semantic deliberately written to avoid giveaway keywords. The remaining 200 questions are where the benchmark earns its keep. Intra-Document Reasoning (40) requires combining distant sections of one long document. Project Related (40) aggregates documents from a single initiative. Constrained (30) supplies several relevant documents but qualifiers that disqualify all but one answer. Conflicting Info (20) contains documents that directly contradict each other, and the README says the system must give a complete and correct answer. Completeness (20) requires fetching every relevant document, capped at 10, to answer correctly.
Two categories have no ground-truth documents at all: High Level (10), where the answer is not located in any single document, and Info Not Found (20), where the answer is not available. Those 30 items are the abstention test. A pipeline that always returns an answer will look fine on Basic and badly here. A further 100 metadata-dependent questions sit in extra_questions.jsonl for metadata-aware RAG, and the README states they are excluded from the leaderboard because their evaluation criteria differ from the core retrieval-focused benchmark.
Installing the evaluation harness and scoring your first run
The repository does not present itself as a pip-installable package. It is a dataset plus generation and evaluation code, and the README points you to two places for instructions: quickstart.md for all the ways to use the dataset and code, and answer_evaluation/README.md if you only want to run answer evaluation. Dependencies are listed in requirements.txt and include openai, anthropic, qdrant-client, opensearch-py, tiktoken, pyarrow and pydantic[email]. Note that qdrant-client and opensearch-py are both present, which tells you the reference setup anticipates more than one vector store backend.
Start by cloning the repository and installing the listed requirements into a virtual environment.
git clone https://github.com/onyx-dot-app/EnterpriseRAG-Bench.git
cd EnterpriseRAG-Bench
pip install -r requirements.txtNext, get the corpus. The README offers two routes: the latest release on GitHub, or the HuggingFace dataset page. Downloading all_documents.zip gives you the whole corpus in a single archive; the per-source slices are named <source_type>_slice_<slice_number>.zip and each holds up to 5000 documents with no nested directory structure. The README describes the archive layout and the slice naming; it does not give an unzip command, so use whatever extraction tool you already have.
After extraction, the question set is in questions.jsonl at the repository root. The README says the question set can be found there and also under the releases. From that point, follow answer_evaluation/README.md for the scoring path rather than improvising your own metric, because the category definitions above assume specific behaviour on Conflicting Info and Info Not Found. If you plan to generate a corpus for your own industry or company scale instead of using Redwood Inference, the README says the code supports that, and methodology.md documents the process; the generation code is under src/.
Where EnterpriseRAG-Bench will mislead you
The benchmark is retrieval-focused, and the README is explicit that the 100 metadata-dependent questions are excluded from the leaderboard precisely because their evaluation criteria differ. If your product's quality depends on metadata filters, recency weighting or permission scoping, the leaderboard number tells you less than you think.
The corpus is also synthetic, and the README frames it as simulating a company. Synthetic noise is not the same as real noise: real internal corpora contain scanned PDFs, broken tables, duplicate attachments and years of abandoned folder structures. A system that scores well here has demonstrated retrieval over generated text, not ingestion resilience. Language is the other boundary. Nothing in the README describes non-English documents or multilingual question answering, so the benchmark is not evidence for a deployment in another language.
One practical constraint: the corpus is large. Slightly over 500,000 documents means embedding and indexing cost, storage and time before you get a single score. Teams that want a fast sanity check on a retriever should start with a small subset of slices rather than the full archive.
Alternatives and how they differ in approach
The obvious comparison is a general question-answering benchmark built on public text, such as the web-search and developer-forum collections the README contrasts itself against. Those give you broad world knowledge and easy access, and they are useful for comparing model reasoning. They will not tell you whether your retriever survives 275,000 Slack messages with conflicting facts, because the failure modes are different: public corpora are clean, deduplicated and written to be found.
A second alternative is to build your own internal evaluation set from your production corpus. That is the highest-fidelity option and the one most teams eventually need, because only your documents contain your codenames. The cost is annotation: you need ground-truth document IDs per question, and you need to cover the awkward categories deliberately. EnterpriseRAG-Bench is a reasonable stand-in before you have that, and the README's generation code exists so you can produce a similar dataset for a different industry or company scale rather than hand-labelling from zero. The trade-off is that a generated corpus inherits the generator's assumptions about what a company looks like.
Licence, maintenance and what an upgrade costs you
The code is MIT licensed, and the repository carries an MIT badge pointing at LICENSE. MIT is permissive, so the practical question is not whether you can use it but what obligations attach to the dataset separately from the code. The README presents the dataset through a GitHub release and a HuggingFace dataset page; if the dataset carries its own terms, they are not stated in the README text available here, and that is worth checking before you redistribute the corpus inside your organisation.
The repository is not archived, and the last push was on 2026-09-03. The only release listed is v1.0.0, tagged Dataset v1.0.0, dated 2026-03-29. That release pattern matters for upgrade cost: dataset releases can change the corpus, and if the documents change, previously computed scores are not comparable across versions. Pin the release you evaluate against and record the tag alongside your numbers. Because the harness is a repository rather than a published package, upgrades mean pulling the branch and re-reading quickstart.md and answer_evaluation/README.md, since there is no versioned API contract to rely on.
Editorial conclusion
Adopt EnterpriseRAG-Bench if you are choosing or tuning a retriever over heterogeneous internal sources and need ground-truth document IDs, and treat the 40 High Level plus 20 Info Not Found questions as your abstention check. Do not adopt it as your only evaluation if your product lives or dies on answer phrasing, or if you need a corpus in a non-English language: the README describes an English-language synthetic company and does not document multilingual generation. Before you commit, open questions.jsonl and confirm that the category counts match the table, and check whether the 100 metadata-dependent questions in extra_questions.jsonl are the ones your pipeline should be graded on, since the leaderboard excludes them.
Frequently asked questions
What is EnterpriseRAG-Bench?
It is a dataset and benchmark from Onyx for retrieval-augmented generation over company internal documents. The corpus contains slightly over 500,000 synthetic documents from nine source types, and the question set contains 500 questions across ten categories.
Where do I download the EnterpriseRAG-Bench dataset?
The README points to the latest GitHub release or the HuggingFace dataset page. You can take all_documents.zip for the whole corpus, or individual <source_type>_slice_<slice_number>.zip files of up to 5000 documents each.
How many documents and questions does EnterpriseRAG-Bench contain?
Slightly over 500,000 documents and 500 questions, plus an additional 100 metadata-dependent questions in extra_questions.jsonl that are excluded from the leaderboard.
Does EnterpriseRAG-Bench evaluate answer quality or only retrieval?
The README describes it as a core retrieval-focused benchmark, and the 100 metadata-dependent questions are excluded from the leaderboard because their evaluation criteria differ. The answer_evaluation directory exists for running answer evaluation separately.
Can I generate a similar dataset for my own company or industry?
Yes. The README states the code provides a way of generating similar datasets for different industries and scales of companies, and methodology.md documents the process behind the five design principles.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/onyx-dot-app-enterpriserag-bench)