# locomo-refined: the judge agreed with humans less than half the time

> LoCoMo-Refined recalibrates a long-conversation memory benchmark on two axes, a stricter judge and a cleaner question set, and reports that the original judge's agreement with human annotators was 43.67% where the refined one reaches 86.33%. The results table needs reading with its definition in hand, because the drop column measures how much the stricter judge corrected each system rather than how any system changed.

**mem-eval-suite/LoCoMo_refined** — LoCoMo Refined: Recalibrating LoCoMo with stricter LLM judging and a cleaned dataset.

- Repository: https://github.com/mem-eval-suite/LoCoMo_refined
- Stars: 672 · Forks: 10
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/mem-eval-suite-locomo-refined

## The judge agreed with humans 43.67% of the time, and now agrees 86.33%

That single pair of numbers is the contribution, and it is worth reading carefully because it is a measurement of the judge rather than of any model.

The at-a-glance table gives the question count, the number of revised samples, the number of human annotators, the number of human-alignment samples used, and then two agreement figures. The original judge's agreement with humans is 43.67%. The refined judge's is 86.33%.

Less than half of the original judge's verdicts matched the humans. Whatever else LoCoMo-Refined does, it starts from the position that the original scoring function was closer to a coin flip than to a measurement, and that any ranking built on it was ranking noise as much as memory.

The dataset itself is 1,382 questions. Five human annotators revised 337 samples, and 300 samples were used for the human-alignment check. So the human work is real but bounded: a few hundred samples out of thirteen hundred, which is the right order of magnitude for an audit and also means most of the question set is unchanged.

The stated motivation matches the numbers. The original evaluation could over-credit answers that were close in topic but wrong in detail, especially around time, missing facts and unsupported additions.

## The drop column measures the benchmark's effect, not the systems' regression

The results table is the part of this page a reader is most likely to get wrong, and one sentence above it is the instruction that prevents it.

The same system predictions were re-scored with the refined judge. The drop is defined as the absolute decrease from the original judge, in percentage points. So it is the size of the correction the stricter judge applied, not a change in the systems.

With that stated, the table reads differently. EverMemOS shows 58.25% with a drop of 22.07 points. MemOS shows 63.60% with a drop of 17.30. MemPalace shows 58.68% with 15.78. Mem0 shows 48.91% with 15.56.

Rank by final score and the order is one thing. Rank by drop and it is another, and the two disagree at the bottom: Mem0 has the smallest correction of the four and the lowest refined score. A small drop does not mean a good score; it means the original judge was already roughly right about that system, which is a different claim.

MemoraX AI is listed with no drop at all, and the reason is in the news section above the table: it reached 82.65% on LoCoMo-Refined in April 2026, so there is no original-judge score for it to fall from. Reading its entry as a smaller drop than the others would be a mistake.

## A correct answer must be inclusive, complete and not overreaching, all three at once

The judge is built around a single sentence, and that sentence is printed twice on the page: once introducing the stricter judger and once as the evaluation principle.

It reads that a prediction is correct if it is inclusive without contradiction and complete without overreach. Expanded into five conditions, a prediction must include all the required information from the gold answer, must not contradict it, must not introduce unsupported extra details, must preserve the correct temporal granularity, and must handle list-style answers without either missing required items or adding unsupported ones.

Every one of those five is aimed at a specific way a memory system can look right while being wrong. Contradiction is a hard error. Missing information is a hard error. Extra detail is an error even though more detail usually reads as a better answer. Wrong granularity is an error even if the date is mentioned. And on list answers, both under-inclusion and over-extension count.

That last group is why the principle is stated as it is. The failure mode being closed is a system that returns a superset of the gold facts and gets credit for effort.

The three behaviours named as what the benchmark is designed to expose are time drift, missing facts, and unsupported claims. Time drift gets its own line in the change table: the original could gloss over vague date conversion or unsupported extra detail, and the refined version requires strict temporal granularity alignment.

## The answer field is a list of alternatives, not a set of required parts

There is one schema decision that will change your score if you get it wrong, and it is spelled out in a parenthetical.

The answer field is a list of acceptable gold answers. Each item in the list is a complete correct answer candidate. And then the clause that matters: if a question requires multiple facts, those facts should appear together inside one answer string, and they are not split across list items as a required set.

So a two-fact question has one entry containing both facts, not two entries. Reading the field as a required set changes what counts as correct and will move your numbers in a direction that has nothing to do with memory.

The data is published in two files. The full set at the raw path holds 1,382 questions. The public questions file is what your prediction identifiers have to match, and those identifiers are conversation-scoped with a question counter inside them.

The refinement covered three named classes of problem in those 337 samples: ambiguous wording, reversed subject and object relationships, and time information that is inconsistent with the original conversations. The reversed relationships are the ones worth thinking about, because a question whose subject and object are swapped is not a noisy question, it is a wrong question.

## Judging is pinned to one model at zero temperature with thinking off

The reference section fixes the judge down to four settings, and each one removes a source of variation.

The model is a specific Qwen checkpoint at fourteen billion parameters. The temperature is exactly zero. Thinking mode is disabled when the model supports disabling it. The prompt lives in one file and the runtime in another.

Zero temperature and disabled thinking together mean the judge is as close to deterministic as the serving stack allows, which matters when the claim being made is a two-figure difference in agreement with humans.

The original judge is kept in the repository and is selectable, so a run can be reproduced on either. The default is the refined one, and the switch is a command line flag rather than a configuration file.

The other guard is louder. If you name a model that is not a Qwen one, the script warns and requires manual confirmation before continuing. The aliases accepted are the hyphenated and underscored spellings of the model name, with or without a vendor prefix, which is a convenience for the two hosting providers that front the same checkpoint.

## Two packages, no manifest, and one shell script with two output formats

The repository root is seven entries and three of them are documents.

There is a licence file and a notice file, plus the README. The other four are the data directory, the scripts directory, the source directory, and nothing else. There is no package manifest, no requirements file and no continuous integration directory.

For a Python evaluation harness that means the two dependencies are named in prose and installed by hand. They are an API client and a retry library, and the setup instructions create a conda environment on a pinned Python version, install both, and then export an environment variable pointing at the interpreter.

That last export is the giveaway that the entry point is a shell script rather than a console command. The script needs to know which Python to run, and it is told through the environment rather than through packaging.

The script itself takes a metrics selection and a judge selection, and it is run twice in the documented flow: once for the lexical metrics and once with the judge added:

```bash
./scripts/run_eval.sh --metrics f1 bleu
./scripts/run_eval.sh --metrics llm f1 bleu --llm-judge refined
```

Three output files come back, one line-oriented, one machine-readable summary and one markdown summary, so the same run serves both a script and a person reading the results.

## A version 1.0.0 release page is linked, and no release is recorded

One of the badges at the top of the page links to a version 1.0.0 release on this repository, and the recorded release list for the repository is empty.

That mismatch is small and worth one sentence rather than a section of analysis. Either a tag was cut without a release object attached, in which case the link resolves to a page describing a tag rather than a release, or the record captured for this repository does not include it.

What matters more for anyone citing this work is the dataset date and the push date. The dataset release is given in the news section as 2026-04-14. The state-of-the-art entry is dated 2026-04-26. The last recorded push is 2026-05-18.

The news list also has an entry that has not arrived yet: a technical report, marked as coming soon. So the reasoning behind a two-figure change in judge agreement is documented in the README as a principle and a table, and the fuller write-up is promised and not yet available.

The recorded state is 672 stars, 10 forks and 11 open issues, with the licence field showing no recognised value even though a licence file sits at the root.

## Conclusion

Use this if you are evaluating memory in a long-conversation system and want a scorer that treats a superset answer as wrong rather than generous, since that is the specific behaviour the recalibration buys. Skip it if you want a fast lexical number, because the lexical metrics ship alongside the judge and the judge is the point. Before you submit a prediction file, read the answer-field rule, because the list is a set of alternative complete answers rather than a set of required parts, and getting that backwards moves your score for reasons that have nothing to do with memory.

## FAQ

### What is the LoCoMo benchmark?

A benchmark for long-conversation memory that asks whether an agent can recall time, events, relationships and preferences after very long dialogues. LoCoMo-Refined is a recalibration of it that tightens the judge and audits the question set, raising the judge's agreement with human annotators from 43.67% to 86.33%.

### Why did the LoCoMo-Refined scores drop?

The same system predictions were re-scored with the stricter judge, and the drop column is the absolute decrease in percentage points from the original judge rather than a change in the system. MemoraX AI is listed with no drop because it was scored on LoCoMo-Refined only, so there is no original score to fall from.

### What model does LoCoMo-Refined use as its judge?

A Qwen checkpoint at fourteen billion parameters, at temperature zero with thinking mode disabled where the model supports it. The prompt is in src/llm_judge.py and the runtime in src/llm_judge_runtime.py. Naming a non-Qwen model produces a warning and requires manual confirmation before the run continues.

### How do I run the LoCoMo-Refined evaluation?

Create a Python 3.11 environment, install the openai and tenacity packages, and export an environment variable pointing at your interpreter. Put your predictions in a JSON Lines file whose question identifiers match the public question file, then run the evaluation script once for the lexical metrics and once with the judge added.

## Sources

- [Issues](https://github.com/mem-eval-suite/LoCoMo_refined/issues)
- [mem-eval-suite/LoCoMo_refined on GitHub](https://github.com/mem-eval-suite/LoCoMo_refined)
- [README](https://github.com/mem-eval-suite/LoCoMo_refined/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/mem-eval-suite-locomo-refined
