gkamradt/needle-in-a-haystack: the niah CLI for long-context retrieval tests
Doing simple retrieval from LLM models at various context lengths to measure accuracy
At a glance
- What is it?
- The v2 rewrite turns the Needle In A Haystack benchmark into a configurable sweep tool: YAML run files, a JSONL result store, and a reconstruct command that rebuilds the exact prompt a model saw. It is a measurement harness, not a library you embed in production.
- Who is it for?
- Adopt it if you need to know at which context length and needle depth a model stops retrieving a planted fact, and you want the exact prompt back when a result surprises you. Skip it if you need a pass/fail gate in CI, a latency benchmark, or a framework to embed in an application: niah writes JSONL and stops there.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 113 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The problem niah measures: retrieval accuracy as context grows
Long-context models are sold on window size, and window size says nothing about whether the model finds a fact buried at token 150,000. The needle-in-a-haystack benchmark exists to answer that narrower question: plant a known string at a known depth in a long document, ask for it, score the answer. The README describes `niah` as running a sweep of `(context length x needle depth)` cells against a configured model, scoring each response, and writing one result row per cell to a JSONL file. The audience is whoever has to choose a model or a context budget and wants numbers rather than a vendor claim: evaluation engineers, applied researchers, and developers comparing two model versions on the same task. The bundled tasks go past single-fact lookup. `single` and `multi` cover one fact and N facts spread through the context, `uuid` asks the model to repeat a fresh UUID, and `uuid_chain` places a chain of `A -> B -> C` links through the document and asks what value is associated with A without revealing the chain structure, so the model has to discover the hops itself. That last task is the interesting one, because it tests multi-step reading rather than string matching.
How a sweep works: run config, model config, recipe
The mechanism is deliberately split in two. A run config names a task, a haystack source, a sweep grid and a store path, and references one model config by id. The model config holds the SDK, the API surface, the environment variable that carries the key, and anything under `request:` is forwarded verbatim to that SDK, so provider-specific knobs such as `thinking` or `reasoning_effort` do not require a code change. The sweep itself is a grid: `context_lengths` and `depth_percents` each take `min`, `max`, `num` and `scale`, plus a list of seeds, so a run is the cross product of lengths, depths and seeds. The runner executes cells with a configurable `concurrency` and `retries`, and `resume: true` lets an interrupted sweep continue against the same store instead of restarting. Results land in JSONL, one row per cell. The design decision worth noting is what is not stored. The README states that rows stay a few KB regardless of context size because the rendered 200k-token context is not saved per row; instead each row carries a recipe with the haystack descriptor, the inserter name, the needle placements (text, `insertion_token_index`, `actual_depth_percent`) and the final token count. `niah reconstruct results.jsonl --row 0` walks that recipe to produce a byte-identical string of what the model saw. That is the right trade-off for a benchmark: storage stays proportional to the number of cells rather than the product of cells and context length, at the cost of a reconstruction step whenever you want to read a prompt.
Installing niah and running a first sweep without an API key
The package is on PyPI as `needlehaystack`, and the console script is `niah`. The README's quick start is two commands, and the first one needs no credentials. It runs a 2 x 3 sweep against an in-process fake model and writes `results.jsonl`.
pip install needlehaystack
niah demo --fakeAfter that, `niah reconstruct results.jsonl --row 0` prints the exact context that row's cell saw, which is the fastest way to confirm the harness did what you expected. To point the same demo at a real model, write a `.env` file with your key and drop the flag. The README gives these examples.
echo "OPENAI_API_KEY=sk-..." > .env
niah demo
niah demo --provider anthropic
niah demo --provider cohereThe default demo uses `gpt-4o-mini`, the bundled Paul Graham essays haystack, a single-fact needle and 6 cells. Once that works, the README says to drop the `demo` command and drive your own sweep with two YAML files, validating before spending money.
niah validate my-run.yaml
niah run my-run.yaml`niah validate` parses and resolves the config without calling a model, so a typo in a model id or a missing environment variable surfaces before the first request. For contributors, the repository ships a fake-provider config that runs the whole pipeline end to end with no keys.
git clone https://github.com/gkamradt/needle-in-a-haystack.git
cd needle-in-a-haystack
uv sync --extra dev
uv run niah run configs/runs/smoke.fake.yamlWhere niah stops being the right tool
This is a measurement harness, and several things follow from that. It does not gate anything: the output is a JSONL file and a CLI, and the README documents no assertion mode, no exit-code-on-threshold, and no CI integration. If you want a build to fail when retrieval accuracy drops, you write that comparison yourself on top of the rows. It is also not a latency or throughput benchmark. Rows carry `duration_seconds` and `usage`, but the runner is designed around a sweep grid with concurrency and retries, not around measuring time to first token under load, so timing numbers from a sweep should not be read as serving performance. The `uuid_chain` task is the one place where scoring is not exact match: the example row shows `score.value` of 0.6 with `details.hops_correct` of 3 out of a chain length of 5, so a chain result is a fraction and you need to decide what fraction counts as success before you compare models. Finally, the project is Python 3.12 and above per `pyproject.toml`, and the v1 `needlehaystack.run_test` entry point is intentionally not declared, so code written against the old API will not import. The README does not document a migration path from v1.
Extending it, and how it differs from a general eval framework
The extension surface is a set of small Protocols connected by registries: `ModelProvider`, `Task`, `HaystackSource` and `Scorer`. A custom task is one class with `generate_needle`, `insert`, `question` and `score`, registered with `register_task` and referenced from a run config as `task.type`. The README states that nothing in the runner needs to change for this, which is the payoff of the registry design.
from needlehaystack.tasks import register_task
class MyCustomTask:
name = "my_task"
inserter_name = "single_depth"
def generate_needle(self, seed): ...
def insert(self, ctx, needle, depth): ...
def question(self, needle): ...
def score(self, response, needle): ...
register_task("my_task", MyCustomTask)The contrast with a general evaluation framework such as promptfoo is the scope of the grid. A general framework asks whether a model passes a suite of prompts, and its unit of work is the test case. `niah` asks how retrieval degrades as two continuous parameters move, and its unit of work is the cell: a context length crossed with a needle depth crossed with a seed. That makes the result a curve rather than a pass rate, and it makes the reconstruction recipe the feature you lean on, because a surprising cell is a specific prompt at a specific depth rather than a named test. A general framework will not hand you back a byte-identical 32k-token prompt from a few KB of metadata. Conversely, `niah` will not run your application's prompts, assert on output shape, or report a suite-level score, so the two are complements rather than substitutes.
Maintenance, licence and the cost of a sweep
The last push to the default branch was on 2026-06-08, and v2.0.0 was released on 2026-05-30 as a clean refactor. The repository is not archived. The version string in `pyproject.toml` is `2.0.0` and the classifiers mark the project as Beta, so the config schema and the JSONL schema should be treated as moving: rows carry a `schema_version` field, currently 2, which is the thing to check when you parse results across releases. `pyproject.toml` declares `license = { file = "LICENSE.txt" }` and the classifier `License :: OSI Approved :: MIT License`, while the repository metadata reports the licence as NOASSERTION. That mismatch is worth resolving against `LICENSE.txt` before you rely on the terms, particularly if you plan to redistribute the bundled Paul Graham essays haystack. Upgrade cost is mostly config drift rather than code: the v1 entry point is gone, so an upgrade from v1 means rewriting run and model files for the v2 schema. Money is the other cost. The model config carries `pricing.input` and `pricing.output` in USD per 1M tokens, and each row records `cost_usd`, so the price of a sweep is knowable before you scale the grid: the README's default demo is described as about $0.01, while the example row for a 32,000-token `uuid_chain` cell records `cost_usd` of 0.171. A grid of 8 lengths by 11 depths by 3 seeds is 264 cells, and that multiplication is the thing to do before running it.
Editorial conclusion
Adopt it if you need to know at which context length and needle depth a model stops retrieving a planted fact, and you want the exact prompt back when a result surprises you. Skip it if you need a pass/fail gate in CI, a latency benchmark, or a framework to embed in an application: niah writes JSONL and stops there. Before committing, run niah validate on your own run config and check the recipe block in results.jsonl, because that recipe is the only record of what the model saw.
Frequently asked questions
What is the needle in a haystack test in this project?
It plants a known string at a known depth inside a long context, asks the model for it, and scores the response, so you can see how retrieval holds up as context length grows. The bundled tasks include single-fact lookup, multi-fact recall, UUID retrieval and UUID-chain hops.
How do I install gkamradt/needle-in-a-haystack?
Install the `needlehaystack` package from PyPI, which provides the `niah` console script, then run `niah demo --fake` to confirm the install without an API key. The fake demo runs a 2 x 3 sweep and writes `results.jsonl`.
Can I run niah without an API key?
Yes. The README documents `niah demo --fake`, which runs against an in-process fake model in about a second, and the repository ships a fake-provider config at `configs/runs/smoke.fake.yaml` for a full end-to-end run.
How do I see the exact prompt a niah result used?
Run `niah reconstruct results.jsonl --row N`. Rows store a recipe instead of the rendered context, and reconstruct walks that recipe to produce a byte-identical string of what the model saw.
Which providers does niah support?
OpenAI, Anthropic and Cohere are supported out of the box, selected with `--provider` on the demo or by the model config on a real run. The README states that adding another provider is a small plugin via `register_provider`.
What are the chances of finding a needle in a haystack in these sweeps?
That is exactly what a run measures: each cell crosses a context length with a needle depth and a seed, and the score column records whether the model retrieved the fact. The `uuid_chain` task reports a fractional score, such as 3 of 5 hops correct.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/gkamradt-needle-in-a-haystack)
Community notes