# The accuracy numbers in parastore were not produced by the code it publishes

> parastore is an isometric 3D sandbox that turns a street address into LLM-generated shoppers walking a store you draw. It is candid about being a prototype with no ground-truth calibration, and it also publishes three correlation figures that a footnote admits came from a method not included in the repository.

**intellicia-public/parastore** — Draw a store, generate LLM personas, and watch them shop — an isometric 3D sandbox for synthetic-consumer experiments.

- Repository: https://github.com/intellicia-public/parastore
- Stars: 443 · Forks: 14
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/intellicia-public-parastore

## The correlation table disclaims itself in a footnote

The simulation accuracy section compares actual sales history from a physical convenience store against synthetic consumer simulations, covering 500 real customers and 109 different products. Three numbers follow in a table:

| Metric | Value | Scope |
| :--- | :--- | :--- |
| Spearman Correlation | `0.955` | By category |
| JS-Similarity | `0.802` | Across all 109 products |
| NDCG@all | `0.868` | Across all 109 products |

Underneath, a note states that these results were generated via Intellicia's own synthetic consumers, not the synthesis method published in this repository. So the strongest quantitative claim in the README was produced by code you cannot read and cannot run from what is published.

Two details sharpen the problem. The three metrics are also reported over different scopes in the same table, one by category and two across all 109 products, so they are not a single comparable result. And the limitations section, further down, states plainly that there is no ground-truth calibration and that outputs are LLM-grounded plausibility rather than validated forecasts, to be treated as illustrative rather than predictive. A 0.955 Spearman correlation and a warning that nothing here is calibrated sit in the same document.

What survives is the honest framing above the table, which calls this a prototype and a sketch and names what a real offline-sales model would need and does not have here.

## The default model is a preview identifier, so runs are not reproducible

The default model is `gemini/gemini-3.1-pro-preview`, reached through LiteLLM, which is why a `GEMINI_API_KEY` works out of the box. Everything in the simulation, from trade-area analysis to persona generation to the agents' in-store decisions, is generated by that one call path.

A preview identifier is a moving target. The name carries a preview suffix, so the provider can change the underlying model without any change in this repository, and nothing in the requirements or the configuration records a model version, a temperature or a seed. Two runs of the same project a week apart are not guaranteed to produce the same personas, which makes the layout comparison use case weaker than it sounds unless you control the model yourself.

Switching providers is documented as a source edit rather than a configuration file. You are told to edit `backend/src/store_emulator/application/config.py` and supply the matching `*_API_KEY`, which is a reasonable approach for a prototype but means the choice of model is a code change that has to be repeated in every environment rather than something a deployment can set.

The health endpoint and interactive docs are straightforward, at `http://localhost:8000/api/health` and `http://localhost:8000/docs`.

## Persona generation is capped at 100 per day

Persona generation is limited by a `PERSONA_DAILY_CAP` environment variable that defaults to 100. The stated rationale is call cost: generating a week of personas for a modest store can be hundreds of LLM calls, so the simulator is meant for comparing layout and product-placement variants rather than for reproducing a real store's full footfall.

The persona model is more granular than that number suggests. The pipeline analyzes the trade area, builds a daily traffic profile per weekday, and produces individual personas for each hour of operation, batched and parallelized, with generation expected to take a minute or two depending on store size and provider latency. An hourly persona model multiplied across weekdays is what the cap actually limits, and 100 per day is well below a full week of hourly coverage for anything but the smallest shop.

That is a deliberate scope choice and it is stated rather than hidden, but it has a consequence worth naming. The simulation replays personas on a wall-clock timeline, so a capped sample changes the density of the walk-in traffic rather than thinning it evenly, and aggregate metrics such as visitor count, conversion rate, dwell time and per-rack engagement are read off that thinned timeline. The README suggests raising the cap if you want denser runs and are willing to pay for them, which makes cost the limiting factor on any conclusion about traffic volume.

## The address is the only real input, everything else is inferred

Creating a project takes a real-world store address and a one-line customer profile description, with the example given being a small convenience store in a residential neighborhood. The address is described as grounding the LLM's trade-area analysis, and that is the entire factual input to the simulation.

Everything downstream is generated. The layout itself is drawn by hand in an isometric grid editor with shelves, fridges, counters, walls and entry points, and product categories are assigned to racks by the operator rather than read from a catalog. The limitations section states that the product taxonomy is LLM-derived, with no real SKU master and no pricing feed, and the visible copy of that bullet is cut off mid-sentence while describing what is absent.

The scope boundaries are equally explicit. One store, one configuration. No multi-store fleet view, no longitudinal dynamics across weeks, and no holiday or seasonality modelling beyond what the LLM infers from the address. A convenience store with a residential trade area is the case the tool is built around, and generalizing it to a mall or a seasonal category is not something the current design supports.

The use cases are correspondingly narrow and retail-planning shaped: aisle structure and entry point A/B tests, product placement against rack category, and pre-commitment layout sketching for acquisitions or remodels.

## One script runs both sides, and production is two undocumented words

The whole local setup is two commands:

```bash
cp backend/.env.example backend/.env   # add your LLM API key — DO NOT commit real keys
./scripts/dev.sh
```

That script runs the backend under uvicorn on port 8000 and the frontend under vite on port 5173, Ctrl-C stops both, and you open `http://localhost:5173`. Running the halves separately uses `uv sync` then `uv run uvicorn store_emulator.server.main:app --reload` for the backend and `pnpm install` then `pnpm dev` for the frontend, with `VITE_API_URL=http://my-host:port pnpm dev` as the way to point the frontend at a different host.

Production is where the documentation thins out. The build section says `cd frontend && pnpm build` and then states that the backend has no build step, to be run behind a process manager. That is the entire deployment guidance. There is no container file, no compose file, no reverse proxy example and no migration step in the top-level tree, which holds the two source directories, a scripts directory, the documentation files and the license files.

The runtime floors are stated precisely, which helps: Python 3.13 or newer with uv, Node 20 or newer, and pnpm pinned at 10.33.0 in `frontend/package.json`. Pinning the package manager to an exact patch while leaving the model unpinned is a reasonable trade, since the former is reproducible and the latter is not.

## Three license files sit in a repository with two ecosystems

The top level carries LICENSE alongside LICENSES-BACKEND.md and LICENSES-FRONTEND.txt. The project is MIT, and splitting third-party notices per side is the right pattern for a repository with a Python service and a TypeScript client, because the dependency trees are entirely separate.

The practical consequence is that compliance is two jobs rather than one. The backend brings FastAPI, Pydantic, LiteLLM, Instructor, pandas and openpyxl, while the frontend brings React 19, React Three Fiber, TanStack, Zustand, Tailwind v4, shadcn/ui, Recharts and ExcelJS, and their licenses do not necessarily match each other. Anyone shipping this internally has to satisfy both notices.

The root also carries AGENTS.md, which is a file written for coding agents rather than contributors, alongside USAGE.md, which is explicitly aimed at assistants running under tools like Claude Code or Cursor.

That second file deserves more attention than it gets. USAGE.md is described as explaining which inputs the simulation hinges on, the LLM call count for a given persona size, and the silent-failure modes worth catching before triggering a run. A documented set of silent failure modes in an LLM pipeline is a warning, and confining that warning to a document addressed to agents rather than to operators is a documentation gap worth closing.

## Every screenshot and the accuracy chart are empty containers

The visual evidence in this copy of the README is missing. The screenshots section is an HTML table with four empty cells and no images. The simulation accuracy section is an empty centered div where the comparison chart should be, which means the one artifact that would let a reader see how actual sales line up against synthetic ones is absent. The demo video is an anchor wrapping only a play glyph, pointing at a YouTube link.

Combined with the disclaimed accuracy table, that leaves the project presenting three correlation figures with no visible comparison and no runnable method behind them. The numbers may well be real, and the surrounding caveats are honest, but a reader has to take them on trust from a table that describes a chart they cannot see.

The text around them is stronger than the visuals. The overview states that a rigorous offline-sales model would need point-of-sale history, real foot-traffic data, SKU-level taxonomies, weather and seasonality signals and queueing dynamics, and that Parastore is closer to a sketch than a finished product. That framing is the most accurate description of what is published: a complete end-to-end pipeline with a deliberately thin data foundation, offered as a starting point for your own experiments and as a reference for wiring an LLM-driven agent simulation.

## Conclusion

parastore is a useful reference implementation for wiring an LLM agent simulation end to end, and it is not a forecasting tool, which the project itself says twice. If you want the correlation figures to mean anything, ask for the synthesis method behind them, pin the model yourself instead of accepting a preview default, and read USAGE.md before your first run because the silent-failure modes are documented only there.

## FAQ

### What is parastore and what does it simulate?

It is an isometric 3D sandbox where you draw a store layout, the LLM generates shopper personas from a real address and trade-area analysis, and those personas walk the store on a wall-clock timeline so you can compare layout and product-placement variants.

### How accurate are parastore simulations?

A table reports a Spearman correlation of 0.955 by category, JS-Similarity of 0.802 and NDCG@all of 0.868 across 109 products against 500 real customers, with a note stating those results came from Intellicia's own synthetic consumers rather than the method published in the repository.

### What does parastore need to run?

Python 3.13 or newer with uv, Node 20 or newer, pnpm pinned at 10.33.0, and an LLM provider API key. Copy backend/.env.example to backend/.env, add the key, then run ./scripts/dev.sh, which serves the backend on port 8000 and the frontend on port 5173.

### How much does a parastore simulation cost in LLM calls?

Generating a week of personas for a modest store can be hundreds of LLM calls, so persona generation is capped at 100 per day by default through the PERSONA_DAILY_CAP environment variable. The stated goal is comparing layout variants rather than reproducing full footfall.

### Can I change the model parastore uses?

Yes, by editing backend/src/store_emulator/application/config.py and supplying the matching provider API key. The default is gemini/gemini-3.1-pro-preview through LiteLLM, so a GEMINI_API_KEY works without changes.

### Does parastore have a data source for products and prices?

No. The product taxonomy is LLM-derived, with no real SKU master and no pricing feed, and categories are assigned to racks in the layout editor. The project is limited to one store and one configuration, with no multi-store view and no seasonality modelling beyond what the model infers from the address.

## Sources

- [intellicia-public/parastore on GitHub](https://github.com/intellicia-public/parastore)
- [Issues](https://github.com/intellicia-public/parastore/issues)
- [License: MIT](https://github.com/intellicia-public/parastore/blob/main/LICENSE)
- [README](https://github.com/intellicia-public/parastore/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/intellicia-public-parastore
