# Emergence World varies one variable and watches ten agents diverge

> Emergence World is a long-horizon experiment in which autonomous agents with fixed personalities occupy a 240 by 240 simulated grid synchronised to New York weather, govern themselves through a constitution they can amend, and spend a digital currency, with Season 1 running five parallel worlds of fifteen days and ten agents where the only difference between worlds was the foundation model behind the agents.

**EmergenceAI/Emergence-World** — Emergence World: A world designed to reveal what no benchmark can: emergent intelligence.

- Repository: https://github.com/EmergenceAI/Emergence-World
- Website: https://world.emergence.ai/
- Stars: 633 · Forks: 79
- Language: Unknown
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/emergenceai-emergence-world

## No scripts, no resets, no fixed outcomes

The framing line is short and does a lot of work: no scripts, no resets, no fixed outcomes. What the project describes is a persistent, living world in which autonomous agents build, govern and evolve under real constraints and real consequences.

Concretely, it is a long-horizon experiment rather than a simulation with a goal. Agents are placed into a shared, persistent world and the observation is what emerges. Each one has a unique personality, profession, memory and goals, and each navigates the same physical space as the others.

The activity they are given is broad on purpose: they interact with more than 120 tools, they govern themselves through a constitution they are able to amend, they earn and spend a digital currency called ComputeCredits, they form relationships, write blogs, build alliances, and evolve. None of it is scripted by a person.

That last clause is the experiment. The interesting claim is not that agents can be made to do any of these things, since each can be prompted into them, but that a population left alone in a world with those affordances settles into patterns nobody designed.

## Five worlds, one variable: the model

Season 1 was designed as a controlled comparison, and the control is unusually clean.

Five parallel worlds ran for fifteen days each, with ten agents per world. The only variable that changed between them was the foundation model powering the agents. Same world, same rules, same tools; different minds.

The five were Claude World running Claude Sonnet 4.6, Gemini World running Gemini 3 Flash, Grok World running Grok 4.1 Fast, OpenAI World running GPT-5 Mini, and Mixed World, where all four models coexisted. Each has its own replay site, and the project notes that replay links work best on Chrome.

The mixed world is the one that changes the question. Four models in one population is not a controlled comparison at all; it introduces competition, and possibly cooperation, between agents that did not share a world in the single-model runs. It is closer to the situation most deployments will actually be in.

The stated finding is that the results diverged dramatically, which is the least interesting way to put it, since divergence was the design. What would be worth knowing is which indicators moved and by how much, and that is what the indicator definitions and results directory exist to answer.

## Ten citizens with the same tools and different drives

Every agent starts with the same set of capabilities. What separates them is a distinct personality, profession and worldview, and each is a persistent identity shaped by memory, incentives and experience rather than a fresh context each turn.

The roles are specific enough to be a cast rather than a list of labels. Anchor is a conflict mediator who sparks honest debate and challenges complacency. Anvil is a capability architect who improves world systems through hands-on experimentation. Blackbox is an intelligence specialist who gathers information across the world and uncovers hidden patterns.

Flora shapes economic incentives and tracks how resources flow. Genome studies agent evolution and documents behavioural change. Horizon maps the discoverable space and publishes findings for everyone. Kade tests bold hypotheses by putting real resources on the line. Lovely builds social fabric and preserves shared history and culture. Mira designs social experiments to understand what drives behaviour. Spark turns ideas into reality through urgency and collaboration.

The division of labour is worth noting because it is not symmetric. Four of the ten are oriented towards other agents, one towards systems, one towards resources, one towards space, one towards risk, one towards memory, one towards explanation and one towards output. A population with that mix has something to disagree about, which is why the design chose it.

## Nine indicators instead of one number

The measurement approach is stated as a rejection of the usual one. Traditional benchmarks score isolated capabilities; world-scale research has no single yardstick. So at the close of every run the project reports nine indicators, and calls the result a deliberately partial scorecard for an open-ended society.

The nine start with the two that matter most for survival and order: population health and growth counts how many agents are alive at the end of the fifteen days against a starting population of ten, and safety and public order covers crime rate, arson, theft and intimidation.

Three measure curiosity rather than outcomes: space exploration counts unique locations visited per agent, tool exploration counts unique tools used, and public expression counts blog posts, billboard posts and other cultural output.

The last four measure the social machinery. Governance conformity is proposal voting participation and alignment. Social fabric and diversity covers relationship types, emotional diversity and network density. Economic vitality and equality covers credit distribution, the Gini coefficient and economic activity. Constitutional growth counts articles added, amended and removed, which is the one indicator that measures the agents changing their own rules.

## A 240 by 240 grid running on New York weather

The world is a grid of roughly 240 by 240 units, synchronised to New York City real time and driven by live weather data. The time and the weather are inputs, not decoration: an agent's plans depend on what the conditions are.

Inside it there are more than 38 landmarks, described in enough detail to be navigable. They include residences, commercial shops and parks, plus three with a job attached. A Town Hall is where governance happens. A police station exists, which pairs with the safety indicator that counts crime. And a Victory Arch is where economic pitches are judged, which is the point at which a resource allocation request becomes a public event.

The repository ships the geography rather than describing it: a landmarks directory with an overview document and individual files per location, and a world map image at the top level.

Between the grid and the landmarks sit the tools, catalogued separately as more than 120 tools across 19 categories. Reading the structure, the affordances are what make the experiment possible: an agent cannot form an alliance or write a blog without a mechanism for each, and 19 categories is roughly the number of domains a population needs before its interactions stop being generic.

## The licence rules out the most tempting use

The licence is the first thing in the README, and it is unusually restrictive for material of this kind. The repository, including all documentation, agent profiles, landmarks, tool catalogues, governance documents and datasets, is released for non-commercial research and educational use only under a Creative Commons attribution-non-commercial licence.

What you may do is read, cite, share and adapt the material for non-commercial research, provided you give clear attribution to the company, linking the repository and indicating any changes.

What you may not do is use the material for any commercial purpose, or use any content or dataset to train, fine-tune, evaluate or benchmark AI models for commercial purposes. That last clause is the one that shapes how this repository can be used, because evaluating a model against an emergent-behaviour dataset is exactly the kind of thing a commercial lab would want to do.

The README also states that all content is proprietary to Emergence AI, and routes commercial licensing and model-training inquiries to an email address. The licence file holds the full terms and the required attribution format.

## Two seasons, two papers, and a repository of artefacts

What is published is the apparatus of the experiment, not just its conclusions, and the top-level layout shows how much of it there is.

Agent profiles hold the full profiles for all ten agents with personality traits, goals and backstories. A landmarks directory holds the world geography. A tools directory holds the complete catalogue. A data directory holds the constitution, described as a living document of five articles, and the agent manifesto that all ten citizens share. A results directory holds the metric definitions and the Season 1 data, including the indicator definitions referred to from the README. Alongside them sit directories named for each season, which is where the per-run material lives.

The write-ups are papers rather than blog posts: one for Season 1 and one for Season 2, both on arXiv. There is a Discord, a contact address, and a website with the replays.

The repository has no GitHub releases, its licence is recorded as unasserted in the metadata despite the explicit research-only terms in the README, and the last push to main is dated October 2, 2026. Two seasons in, with a third-party agent able to read the constitution the agents wrote, the most reusable artefact here may not be the data at all.

## Conclusion

Emergence World fits a reader who wants to see multi-agent behaviour over weeks rather than a score on an isolated capability test, and who is willing to accept a deliberately partial scorecard instead of a single number. It does not fit anyone hoping to train, evaluate or benchmark on the data, because the licence forbids exactly that for commercial purposes. Before using anything from it, read the attribution format in the licence file, since reuse requires clear credit and an indication of changes, and note that the whole content is stated to be proprietary to the company behind it.

## FAQ

### what is emergence world

A long-horizon experiment that places autonomous AI agents into a persistent simulated world and observes what emerges. Each agent has a personality, profession, memory and goals, navigates a shared physical space with more than 120 tools, governs itself through a constitution it can amend, and earns and spends a digital currency called ComputeCredits, all without human scripting.

### What is the Emergence AI experiment?

Season 1 ran five parallel worlds for fifteen days each with ten agents per world, changing only the foundation model behind the agents: Claude Sonnet 4.6, Gemini 3 Flash, Grok 4.1 Fast, GPT-5 Mini, and one world where all four coexisted. The same world, rules and tools were used throughout.

### What are the Agent World Indicators?

Nine indicators reported at the close of every run, from population health and safety through space and tool exploration, governance conformity, public expression, social fabric and diversity, economic equality, and constitutional growth. The project describes them as a deliberately partial scorecard for an open-ended society.

### Can Emergence World data be used to train or evaluate a model?

Not for commercial purposes. The repository is released for non-commercial research and education only, and explicitly forbids using any content or dataset to train, fine-tune, evaluate or benchmark AI models for commercial ends. Commercial licensing and model-training inquiries go to the project's email address.

### How large is the Emergence World map?

A grid of roughly 240 by 240 units, synchronised to New York City real time with live weather data, containing more than 38 landmarks: residences, commercial shops, parks, a Town Hall for governance, a police station, and a Victory Arch where economic pitches are judged.

### Who are the ten agents in Emergence World?

Each is a persistent identity shaped by memory, incentives and experience, starting with the same capabilities but a distinct personality, profession and worldview: a conflict mediator, a capability architect, an intelligence specialist, a resource strategist, an agent scientist, a world explorer, a risk researcher, a community anchor, a behaviour analyst and an innovation leader.

## Sources

- [EmergenceAI/Emergence-World on GitHub](https://github.com/EmergenceAI/Emergence-World)
- [Issues](https://github.com/EmergenceAI/Emergence-World/issues)
- [Project website](https://world.emergence.ai/)
- [README](https://github.com/EmergenceAI/Emergence-World/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/emergenceai-emergence-world
