# JPAF's headline accuracies are one model judging another, and the citation is commented out

> A research framework for evolving agent personalities, built around three named mechanisms and three evaluation methods, shipped as two experiment directories, one requirements file and a per-directory secrets file. No licence, no releases, and no commits since March 2026.

**agent-topia/evolving_personality** — Dynamic MBTI Personality Simulation for LLM Agents via Carl Jung's Theory. A framework that enables LLM agents' MBTI personalities to naturally evolve and grow through interaction.

- Repository: https://github.com/agent-topia/evolving_personality
- Website: https://arxiv.org/abs/2601.10025
- Stars: 1,149 · Forks: 104
- Language: Python
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/agent-topia-evolving-personality

## Every accuracy number comes from one model judging another model's answers

The experimental highlights are four figures: 100 percent MBTI alignment accuracy across all tested models, type activation accuracy above 90 percent for GPT and Qwen and between 65 and 95 percent for Llama, and personality evolution accuracy of 100 percent for GPT and Qwen against 92 percent for Llama. The mechanism behind them is in the script names rather than in the prose. There is a judge method, and the documentation says you run it to obtain judge results in different models. Every sample answer in the output carries a free-text reason field, which is exactly what a judge model would read. So the measurement is one model scoring another, with no sample size, no protocol description and no second baseline given for any of the four figures.

## The judge run is documented with different parameters from the runs it judges

Three commands are given, one per method, and they do not match each other. The judging step is shown processing one type and one sample:

```py
python personality_test.py \
    --method=judge \
    --mbti_num=1 \
    --model=QWEN \
    --nums=1 \
```

The baseline and the framework run are both shown with sixteen types, five samples and seventy test items:

```py
python personality_test.py \
    --method=test \
    --mbti_num=16 \
    --model=QWEN \
    --nums=5 \
    --test_num=70 \
```

So the stage that produces the scores is documented at a scale of one, while the stages being scored run at sixteen and five. The baseline is also worth noting for what it is: the no-prompt method measures an unconditioned model, not another personality framework.

## The README says eight types in one paragraph and sixteen in another

The key features section describes the modelling as based on Jung's eight psychological types, with fine-grained personality expression through weighted type differentiation. Three paragraphs later, the highlights section says the framework supports dynamic simulation and evolution of all 16 MBTI personality types. Both counts appear in the same document and neither is reconciled with the other. Between them sit the three mechanisms, which are the actual design: dominant-auxiliary coordination to hold the core personality steady, reinforcement-compensation for short-term adaptation to context, and a reflection mechanism to drive longer-term change. That progression is the most reusable thing here, because it maps three distinct timescales onto three separate mechanisms rather than one prompt that is asked to do all three.

## The contents list hides its citation, its roadmap and a question about humanoid agents

The table of contents names four live sections, covering installation, customising the model API, the verification experiment and the change experiment. Below them sits a block commented out in the markup, listing four more that do not exist: an analytics dashboard, a section explaining how humanoid agents work, future plans, and a citation. For a project whose repository homepage is an arXiv identifier and whose single news item announces that its paper has been uploaded, having the citation section commented out is the most consequential omission on this list. The commented humanoid question matches nothing else in the repository, whose only directories are two experiment folders, an assets folder and two readme files.

## Conda is a hard prerequisite for what is a small Python codebase

The prerequisites list three things: an operating system, which can be macOS, Linux or Windows, Python 3.10 or newer, and a package manager, named as conda. The installation then creates a conda environment pinned to Python 3.10, activates it, installs from a single requirements file whose name is singular, and clones with the branch named explicitly. For someone reproducing a paper that is entirely defensible. For someone who wants to read the code, it is heavier than it needs to be, since the repository has no package manifest, no installable module and no console entry point, just two directories of scripts and a requirements file. The clone block is short enough to read in place.

## The same credentials get copied into two files, and only one provider has a real endpoint

Configuration is a file named para.env rather than anything the usual tooling looks for, and the installation step copies the example into two separate locations, once for the verification experiment and once for the change experiment. That means the same set of API keys lives in two files by design, and there is no note about which one wins if you edit only one. The example declares a single model selector with three allowed values and three provider blocks. Only the OpenAI block ships a working default endpoint; the Qwen and Llama base URLs are placeholders you have to fill. The two model defaults are large hosted ones, a 235B-class Qwen instruct model and a fourth-generation Llama Maverick, which is a substantial bill for a personality experiment unless you point them somewhere else.

## No licence, no releases, and the last commit is seven months old

Three facts about the repository as an artefact. The repository names no licence, and there is no licence file among the top-level entries, which matters for a framework that also exists as a paper and that anyone might want to build on. There are no GitHub releases, so there is no version to cite or pin, and the only version signal is the commit history. And the last push is dated 2026-03-17, two months after the news item announcing the paper, and about seven months before the current date. The paper's identifier is dated January 2026, so the repository and the paper are both recent and both quiet.

## Conclusion

The three mechanisms are a clean idea and worth reading on their own, because they separate the three things personality work usually conflates: staying recognisable, adapting to the moment, and changing over time. The three evaluation methods are also the right shape, separating an unconditioned baseline from a judge from the framework's own run. What is missing is the part that would let anyone else reproduce it. The accuracy figures come from a judge model reading free-text explanations, with no sample size, protocol or baseline named. The parameters in the documentation for the judging stage do not match the runs it is judging. The repository has no licence and no citation section. Treat this as a paper you read and a codebase you inspect, not as a framework you adopt, and settle the licensing question before anything of yours depends on it.

## FAQ

### What is the JPAF personality framework?

A Jungian-psychology-based framework that gives language models structured personalities through three mechanisms: dominant-auxiliary coordination for core consistency, reinforcement-compensation for short-term adaptation, and a reflection mechanism for long-term evolution.

### How do I install the JPAF personality framework?

Clone the main branch, create a conda environment on Python 3.10, install from the single requirements file, then copy the example configuration into both the Personality_test and Personality_changes directories and fill in your provider credentials.

### Which models does JPAF support?

Three provider blocks: OpenAI with gpt-4 as the default, Qwen with a 235B-class instruct model, and Llama with a fourth-generation Maverick. The project claims validation on GPT, Llama and Qwen model families.

### What accuracy figures does JPAF report?

100 percent MBTI alignment across tested models, type activation above 90 percent for GPT and Qwen and between 65 and 95 percent for Llama, and evolution accuracy of 100 percent and 92 percent. No sample size or protocol is given, and the scoring stage is a model reading another model's answers.

### How does JPAF evaluate personality?

With three methods chosen by a flag: a judge method that has one model score answers, a no_prompt method that measures an unconditioned model as a baseline, and a test method that runs the framework itself.

### What license is JPAF released under?

The repository names no licence and contains no licence file, which is worth resolving with the authors before building on it or redistributing it.

## Sources

- [agent-topia/evolving_personality on GitHub](https://github.com/agent-topia/evolving_personality)
- [Issues](https://github.com/agent-topia/evolving_personality/issues)
- [Project website](https://arxiv.org/abs/2601.10025)
- [README](https://github.com/agent-topia/evolving_personality/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/agent-topia-evolving-personality
