Model or dataset
nokia-applied-research/AnyJev avatar
nokia-applied-research/AnyJev

AnyJev: turning a next-token distribution into a decision you can threshold

Turn any LLM into a Jev-style decision model: typed decisions, real probabilities, no training. (continue updating, welcome any issue and PR request)

1,146 stars138 forksPythonApache-2.0

At a glance

What is it?
A Nokia and Tencent research project that reads a probability out of one prefill instead of parsing generated text, with four cost levels from zero labels to a closed-form head. The packaging does not install the server its own quickstart needs.
Who is it for?
AnyJev is worth understanding if you are routing support tickets, triaging intake or doing any classification where a confidence value has to mean something before you automate on it. The order-flip problem is real and the zero-label fix is the part that needs no budget argument, while the calibration numbers are what actually move the share of traffic you can send without a human.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One prefill, no generation, no parsing

The premise is narrow and specific. Ask an open LLM a typed question, a choice, a yes/no or a score, and get back a decision with a probability attached, read from a single prefill of the next-token distribution. Nothing is generated, so there is no regex to maintain, no JSON schema to retry on, and no partial output to reconcile when a model decides to add a preamble. The README's framing of the problem is that raw logits change their answer when you reorder the options, and their confidence cannot be trusted.

Both halves of that are claims you can check. The order sensitivity is demonstrated with a GIF on a real BANKING77 item with Qwen3-8B, where the raw logit readout flips under option reversal while the L0 readout does not. The caption states that every number in it is a model output, which is the sort of claim most projects do not make about an animation.

The project comes from Nokia's Sunnyvale research group with an author at Tencent Hunyuan, is Apache-2.0 licensed, has 1,081 stars and 136 forks, and was pushed on 2026-10-02 with 13 open issues. Both backends are first class rather than bolted on: `HFBackend` reads hidden states from transformers, and `VLLMBackend` talks to a served model over HTTP. Repository topics list `calibration`, `jev`, `system-one`, `transformers` and `vllm`, which is an unusually honest topic list for a project that supports both.

Four levels, and the price of each one

The results table compares three configurations on Qwen3-8B over BANKING77 as a 20-way choice with 300 test items. Raw logits need no labels, flip their answer on option reversal at a rate of 0.230, reach 0.747 accuracy, carry calibration error of 0.240, and leave only 7.7% of traffic auto-decidable at 5% error. L0 needs no labels either, drops the flip rate to 0.073, lifts accuracy to 0.803 and calibration error to 0.184, and reaches 46.3% auto-decidable. L1 adds a temperature and 100 to 500 labels, giving 0.077 flips, 0.807 accuracy, 0.095 calibration error and 52.0% auto-decidable.

The argument the README makes is about the last row rather than accuracy. Moving from raw logits to L1 is about six points of accuracy, but the share of decisions you can act on without review goes from 7.7% to 52.0%, because with raw logits a stated confidence of 0.9 is not trustworthy enough to act on and everything goes to a human. L2, described in its own section, adds a closed-form head per question solved on 100 to 300 labels in seconds, with no gradients and no training loop.

One detail in the table deserves attention because it is easy to skim past. L0 fixes the order-flip problem without labels, which is a genuine architectural result, but it does not by itself fix calibration, and its ECE of 0.184 is barely better than raw. So if your problem is an unstable answer under option ordering, L0 is the fix. If your problem is a confidence number you cannot trust, you are buying L1 or L2 with labels.

The rotation budget certifies itself against its own full readout

The most interesting mechanism is the one that needs no labels at all. For a K-option choice, L0 can ask the model once per option rotation so that no option is favoured by position. Doing all K rotations is wasteful, so the library can read them one at a time and stop when the leader is far enough ahead. The threshold is not guessed: `Decider.calibrate_adaptive(question, states, target=0.01)` reads every rotation of a batch of unlabelled states once and returns the cheapest threshold whose disagreement with the full-K answer is under the target by a Clopper-Pearson upper bound.

What makes this defensible is the choice of reference. The full-strength readout is the library's own, never a human label, so the guarantee costs calibration compute rather than annotation. Before calibration the threshold is a documented constant of 8.5, described as the smallest value that certified 1% on four model and task cells at once. The reported outcome on Qwen2.5-7B and Qwen3-8B over two tasks is 5.1 to 9.7 rotations read uncalibrated, against 7.2 instead of 18 on the headline example, at 2.2 times the decisions per second on vLLM and 2.3 to 2.7 times on transformers with accuracy unchanged.

The diagnostics are the part to look for when you use it: `decide_batch` reports `shifts_used` and `stop_threshold`, and L2 decisions report `blocks_executed`. Being able to see what the stopping rule did per call is what makes a probabilistic shortcut auditable rather than magic.

Truncating the model made results slightly better

The README states that depth is usually a gain rather than a trade, and then gives the measurement. Cutting Qwen2.5-7B from 28 blocks to 18 left accuracy slightly higher and calibration better, and was faster. The explanation given is that a middle block is a better feature space for a linear head than the last one, where the remaining blocks are busy turning the answer into tokens.

That is a claim most inference optimization projects would avoid making, because the usual expectation is that removing layers must cost something. It follows from the same design choice that makes L2 cheap: if the head reads a hidden state rather than a logit, the question becomes which hidden state, not which token. The truncation step is a separate command, and the comment beside it in the quickstart notes that you keep the blocks a decision needs, usually about two thirds:

bash
python -m anyjev.truncate Qwen/Qwen2.5-7B-Instruct 18 ./qwen-b18

Quantization is offered and discouraged in the same breath. `--quantization fp8` is available and not recommended, because it buys single-question latency and costs accuracy. For a library whose selling point is a trustworthy probability, that is the right default advice, and it is also a reminder that the latency numbers elsewhere in the README were measured at full precision.

The serving quickstart needs a server the package never installs

Here is the practical gap. The README opens with a section titled Serve it and gives three commands as the main path. Step one installs the project with the `hf` extra, step two truncates a model, and step three serves it:

bash
vllm serve ./qwen-b18 --task embed \
  --override-pooler-config '{"pooling_type":"LAST","normalize":false,"softmax":false}'

The `pyproject.toml` does not mention vLLM anywhere. Hard dependencies are numpy 1.24 or newer and nothing else. The `hf` extra is torch and transformers, the `bench` extra is datasets and scikit-learn, and `dev` is pytest and ruff. So the headline workflow depends on a binary that the package assumes you already have, and the two Python imports in the example, from `anyjev` and from `anyjev.backends.vllm`, only work once it is on the path.

The design behind it is still the strongest argument in the README. Because L2 reads a hidden state rather than a logit, an L2 deployment is described as a pooling server plus a few kilobytes of head, with no logits, no parsing, no patched engine and nothing generated. The README also reports that a head fit through transformers and served by vLLM matches one fit and served on either alone, at 99.0% identical answers and a mean absolute probability difference of 0.0011 on BANKING77-20, which is what makes the backend choice a deployment detail rather than a modelling commitment.

For your own numbers, the repository ships a single command that truncates, serves, fits, measures accuracy, ECE and milliseconds per decision on held-out states, shuts the server down, repeats at full depth for comparison, and prints the spread next to the median because a shared machine can report the same configuration as both faster and slower:

bash
python -m anyjev.pipeline Qwen/Qwen2.5-7B-Instruct --labels-from banking20

A Pre-Alpha classifier sitting next to a full results table

Two packaging details set expectations the README does not. The first is the trove classifier, which reads `Development Status :: 2 - Pre-Alpha`. That sits oddly beside a results table with per-level ablations, a linked evidence document for every number, a reproducibility command and three published releases. The second is the description string, which ends with a disclaimer that the project is not affiliated with TypeSafe AI, while the tree carries `CREDITS.md` and `THIRD_PARTY.md`. If you are evaluating provenance rather than capability, those two files are where the answers are.

The repository structure is worth a glance because it separates concerns unusually well. `anyjev/` holds the package, `anyjev-heads/` holds fitted head artifacts, `bench/` the benchmark harness, `demo/` a small application with a games subdirectory and recorded results, `scripts/` reproducibility helpers including an exit-parity check and a script that hunts for genuine option-order flips, and `space/` a Hugging Face Space entry point. Tests live in `tests/` and pytest is configured against that path with a 120 character ruff line length at py310.

One version detail deserves a decision rather than a shrug. The 0.2.0 release notes say `adaptive_shifts=True` is now the default, while the README still frames rotation budgeting as something you turn on and recommends it for any K-option choice, adding that it is opt-in in 0.2.0 because the documentation tables were measured before it existed and that it becomes the default in the release that regenerates them. Both statements are in the project's own words. Until the tables catch up, pass the flag explicitly so your behaviour does not depend on which reading is current.

Editorial conclusion

AnyJev is worth understanding if you are routing support tickets, triaging intake or doing any classification where a confidence value has to mean something before you automate on it. The order-flip problem is real and the zero-label fix is the part that needs no budget argument, while the calibration numbers are what actually move the share of traffic you can send without a human. Two things to check before committing. The package declares only numpy as a hard dependency and never declares vLLM anywhere, so the documented serving path needs a manual install. And the rotation budget defaults are described differently in the 0.2.0 release notes and the README, so pin the flag explicitly. Run `anyjev.pipeline` on your own box before believing any number in the table, including theirs.

Frequently asked questions

Do I need labelled data to use AnyJev?

Not for the first two levels. L0 reads a probability from option rotations and needs no labels at all, which is what fixes the order-reversal flips. L1 adds a temperature and takes 100 to 500 labels, and L2 fits a closed-form head on 100 to 300 labels in seconds with no gradients. The zero-label path fixes answer stability but only partly improves calibration error, so pick the level from the problem you have.

What does AnyJev do differently from asking a model for JSON?

Nothing is generated. The decision comes from one prefill of the model's next-token distribution, so there is no output to parse, no schema to retry and no partial answer to reconcile. That also means no cost from decoding tokens, and the deployable artifact is a small head rather than a prompt pipeline. The trade is that the question must be typed as a choice, yes/no or score rather than open-ended.

Does AnyJev need vLLM to run?

Not strictly, because a transformers backend is included, but the documented serving path does. The pyproject declares only numpy as a hard dependency, with torch and transformers behind the hf extra and no mention of vLLM in any extra, yet the README quickstart serves the truncated model with the vllm command before importing the VLLM backend. Install vLLM yourself before following that path.

Can I reproduce the numbers in the AnyJev README?

Yes, on your own hardware. Running the pipeline command for a model with a labels-from option truncates it, serves it, fits a head, measures accuracy, calibration error and milliseconds per decision on held-out states, shuts the server down and repeats at full depth for comparison. Reported timings are a median over several passes with the spread printed alongside, because a shared machine can make the same configuration look both faster and slower.

Official sources

  1. License: Apache-2.0
  2. nokia-applied-research/AnyJev on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nokia-applied-research-anyjev.svg)](https://hysenlabs.com/projects/nokia-applied-research-anyjev)