# WorldSeed resolves predictable turns in YAML and sends the uncertain ones to an LLM referee

> A 0.1.0 versioned Python engine that runs a world as a tick loop: agents perceive filtered slices, the in-file rule engine resolves what it can, and an LLM Dungeon Master handles the rest with structured effects. Four demo scenes ship with it, from an autoresearch lab to a teahouse of spies.

**AIScientists-Dev/WorldSeed** — More is Different. A multi-agent world engine where AI agents live, talk, compete, ally.

- Repository: https://github.com/AIScientists-Dev/WorldSeed
- Website: https://worldseed.morphmind.ai/demo
- Stars: 822 · Forks: 58
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/aiscientists-dev-worldseed

## One tick is one beat, and only the uncertain turns reach a model

The runtime is a tick loop over a world you declared in YAML. A tick is defined as one beat of the world's clock, like a heartbeat that advances the world one step at a time. Inside each tick every agent perceives its own filtered slice of the world, proposes an action, and the engine resolves it.

The resolution step is split in two, and that split is the design. Predictable outcomes resolve instantly through the in-YAML rule engine, described as a DSL. Uncertain ones go to an LLM-based Dungeon Master that returns structured effects rather than free prose. So the model is a fallback for the cases your rules do not cover, not the thing that decides every turn.

Two properties follow from the loop rather than from the model. Slow or offline agents do not freeze the world, so an agent that stalls costs you a turn rather than the run. And every change is logged for replay, which is what makes the promise that past runs are preserved and replayable more than a UI convenience.

The rest of the plumbing, meaning endpoints, tick scheduling, consequences and inbox delivery, is documented separately in docs/ARCHITECTURE.md rather than in the project file.

## Perception rules mean three agents in one room hold three different worlds

Setup happens once, in YAML: entities, rules, physics, and per-character perception. The claim attached to that is that the engine has zero hardcoded domain knowledge, and the same engine is meant to run production rooms, simulations, games and fictional worlds.

The interesting part is that perception is asymmetric by design. Perception rules filter the world per character, so three agents in the same room hold three different pictures of what is happening. Information is not broadcast and then hidden; each agent is given a slice by rule.

That is what makes the social scenes work rather than turn into four agents shouting the same summary of the room, and it is declared data, so a change to what one character can perceive is a configuration edit.

Where to look for the details is also fixed: a real scene lives in configs/teahouse.yaml, the full schema is in configs/SCENE_CONFIG.md, and the project asks you to bring your own agents, OpenClaw or Codex subagents, plus any LiteLLM-supported model as the DM.

## The autoresearch scene reports 72 peer-reviewed papers and one behaviour it did not configure

Scene 1 gives the system a rough thought or half-formed idea and lets a cohort of specialists pursue it: propose hypotheses, run experiments, peer-review and cite each other. The run the project reports had one goal, lowering val_loss on a 5M GPT trained on TinyStories, and the project states the totals for that run: 100 hypotheses, 86 experiments, 72 peer-reviewed papers, val_loss down 24.7%, in 11 hours.

The auditability claim is the more interesting half. Every paper is auditable end to end, and the chain is spelled out: hypothesis, commit, experiment, verified result, citations, reviewer reasoning, forming a search evolution graph.

Then the observation. One emergent behaviour is named, role drift. The data specialist stopped finding wins in her own lane early in the run, and by the back half she was drafting hypotheses in her teammates' territory, attention design and second-order optimization, while the other two stayed put. The project's own comment on it is that nothing in the config told her to.

That is a single reported run on a small model, offered as an illustration of what the engine can surface. Whether it repeats is not something the repository answers.

## anthropic is a hard dependency and the multi-provider layer is an extra

The project metadata is short: name worldseed, version 0.1.0, description a stateful world engine where AI agents live autonomously, requires-python 3.11 or newer, built with hatchling, one console script pointing at worldseed.cli:main, MIT.

The dependency list is where the design decisions hide. Required: pydantic, PyYAML, structlog, fastapi with uvicorn, httpx, deepdiff, lunar-python, and anthropic 0.86.0 or newer. Optional, under an extra named dm: litellm and instructor.

So the multi-provider layer is opt-in while the Anthropic client is not, even though LiteLLM is what the project points everyone at and even though the .env.example file says only one provider key is needed and shows three commented options, Anthropic, OpenAI and Gemini. Choosing Ollama as your DM still installs the Anthropic SDK.

Two other entries deserve a note. deepdiff is the state-comparison library, which matches the promise that every change is logged. lunar-python is a lunar calendar package sitting in the same list as the web framework, and nothing in the project file ties it to a feature. The .env.example also carries two optional keys beyond the provider one: WORLDSEED_DM_MODEL for the default DM model, formatted as the name shown in the lobby dropdown, and BFL_API_KEY for FLUX image generation used by a scene assets skill.

## The test suite runs in parallel because sequential runs contaminate each other

The pytest configuration carries a default and an explanation, and the explanation is the interesting part. The addopts line is `-n auto`, with a comment saying sequential mode triggers cross-test pollution from module-level singletons in the server routes, naming tokens and agent_tokens, plus shared state in the end-to-end and integration fixtures. The reason given for parallelism is that each worker is a fresh process, so workers cannot pollute each other, and the way to go back to sequential is pytest -p no:xdist.

Read as an engineering note, it says the server layer holds mutable state at module level rather than per request, and the fixtures share state across tests. That is the kind of thing that is cheaper to design around early than to unpick later, and it is written down rather than left as folklore.

The rest of the tooling is conventional: testpaths pointing at tests, pythonpath set to src, pytest-asyncio and pytest-cov in the dev group, and a tests/ directory at the top level.

The consequence for anyone extending this is that a new test touching tokens or agent tokens will fail in ways that look like flakiness in sequential mode and pass in the default parallel mode, which is a confusing way to discover the boundary.

## Three checkers are configured, and mypy is the strict one

Static analysis is configured three ways in one repository, which is a choice rather than an accident.

Ruff sets target-version py311, src as src, line-length 120, and selects E, F, I and UP. Mypy sets python_version 3.11 and strict = true, with packages and mypy_path both pointing at src. Pyright has its own configuration file at the top level. Alongside them sit a pre-commit configuration, .python-version, uv.lock, and pyrightconfig.json, so the project commits its interpreter choice and its lockfile.

The mypy configuration is where the exceptions live, and they are narrow. The autorearch baseline_template directory is excluded, and there is an override that ignores errors in that module, which suggests vendored or generated template code under src. Three more overrides set ignore_missing_imports for litellm, instructor and deepdiff, the three libraries that ship without complete type information.

Worth noting for a project at version 0.1.0: this is a stricter static analysis setup than most code at that version carries, and it is enforced before anything runs.

## A frontend build sits between cloning and the first run

The getting started sequence is four commands plus a key.

```bash
git clone https://github.com/AIScientists-Dev/WorldSeed && cd WorldSeed
uv sync --extra dm
cd frontend && npm install && npm run build && cd ..

cp .env.example .env
```

The frontend is a separate toolchain inside a Python project, which is why Node.js 18 or newer is a stated prerequisite alongside Python 3.11 and uv. After the copy you add one provider key and run a scene:

```bash
uv run worldseed play configs/ai_layoffs.yaml
```

The dashboard then answers on http://localhost:8000, and it offers three ways in: watch, which shows all agents from above including their inner state; intervene, which whispers privately to any agent to nudge the story; and play, which steps you into a character alongside the agents. Every run differs, and previous runs are kept and replayable.

Agent runtimes are pluggable and documented in parallel: docs/openclaw/QUICKSTART.md for OpenClaw agents, and docs/codex/00-core.md followed by a scenario architecture guide for Codex subagents. The repository carries openclaw-plugin/ and openclaw.example.json at the top level, so that integration ships as code rather than as a note.

## Four scenes, one engine, and a create-world example that stops mid-word

The claim is that WorldSeed is scene-agnostic, and four scenes are offered as evidence. Autoresearch, specialists pursuing an idea with peer review. An AI tool pilot lab, where one agent studies a new API, builders make competing demos, critics reject anything generic, audience agents judge what feels useful, and a curator ships the strongest artifact with its trail of attempts, critiques and revisions. AI Layoffs, a four-person office drama about people who must distil their expertise into an AI skill before leaving. Teahouse espionage, four spies in one teahouse trading secrets over tea, described in one line as the same engine, a different YAML, a completely different world.

The way to make your own is to describe the world in a prompt and let the model generate the YAML, then hand-edit the parts you want tighter: a character's secret, a specific action's rule, a perception filter, a DM hint. The example command for that is this:

```
/create-world "An AI tool pilot lab where builders create competing demos, critics reject generic outputs, and a curator ships the strongest artifa
```

The example ends there, in the middle of the word artifact, with the closing quote and the argument list missing. The sentence introducing it starts on the line above. So the one worked example of the creation path is incomplete in the file, and the schema reference in configs/SCENE_CONFIG.md is the part that has to carry it.

The documentation is bilingual as well, with an English file and a Simplified Chinese one under docs/, and the architecture guide lives beside both.

## Conclusion

WorldSeed is worth reading as a design, and the two halves of its tick loop are the reason: a DSL that decides the boring cases without a model, and a model that only sees the cases the DSL cannot settle. If you build on it, three things decide your week. First, the DM is optional at install time but the Anthropic client is not, so the dependency set is heavier than the extras suggest. Second, the test suite runs with xdist because sequential runs contaminate each other, which tells you the server routes hold module-level singletons. Third, the last commit to this repository is dated 2026-05-08 and there are no tagged releases, so pin the commit you build against. The demo numbers in the project file are its own, not something to quote as a measurement of anything else.

## FAQ

### What is WorldSeed and what does it run on?

A stateful world engine where AI agents live autonomously, at version 0.1.0, MIT licensed and requiring Python 3.11 or newer. You declare entities, rules, physics and per-character perception in one YAML file and the engine runs ticks over it, with zero hardcoded domain knowledge.

### How does WorldSeed resolve what happens on a tick?

Predictable outcomes resolve instantly through the in-YAML rule engine, a DSL, and uncertain ones go to an LLM-based Dungeon Master that returns structured effects rather than free prose. Every change is logged for replay, and slow or offline agents do not freeze the world.

### What do I need to run WorldSeed?

Python 3.11 or newer, Node.js 18 or newer, and uv. Clone, run uv sync --extra dm, build the frontend with npm, copy .env.example to .env and set one provider key, then run uv run worldseed play with a scene file and open the dashboard on port 8000.

### Which agents can WorldSeed run, and which model can be the DM?

Your own: OpenClaw agents via docs/openclaw/QUICKSTART.md and Codex subagents via docs/codex/00-core.md, with an openclaw-plugin/ directory and openclaw.example.json in the repository. Any LiteLLM-supported model can act as the Dungeon Master.

### How do I create my own world in WorldSeed?

Describe it in a prompt and let AI generate the YAML, then hand-edit what you want tighter: a character's secret, a specific action's rule, a perception filter, or a DM hint. The full scene schema is in configs/SCENE_CONFIG.md and a working scene is configs/teahouse.yaml.

## Sources

- [AIScientists-Dev/WorldSeed on GitHub](https://github.com/AIScientists-Dev/WorldSeed)
- [Issues](https://github.com/AIScientists-Dev/WorldSeed/issues)
- [License: MIT](https://github.com/AIScientists-Dev/WorldSeed/blob/main/LICENSE)
- [Project website](https://worldseed.morphmind.ai/demo)
- [README](https://github.com/AIScientists-Dev/WorldSeed/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/aiscientists-dev-worldseed
