MingLi-Bench grades exact match against a fortune tellers competition sheet, and its two recommended flags are off by default
A benchmark for evaluating LLMs on Chinese traditional fortune telling — Bazi (八字) and Ziwei Doushu (紫微斗数).
At a glance
- What is it?
- A Python benchmark that asks models 160 multiple-choice questions drawn from an annual Chinese fortune telling competition, scores by exact match, and can inject pre-computed Bazi and Ziwei charts so the number reflects reasoning rather than chart arithmetic. No releases, last commit 2026-05-09.
- Who is it for?
- MingLi-Bench is a measuring stick, not a model, and it is worth running if your interest is how language models handle Chinese technical prose under exact-match scoring. Two things decide whether the number means anything: whether you pass --cot and --astro, because the recommended configuration is not the default one, and whether the keys in your .env point where you think they do, because the dotenv block in the README and the .env.example in the tree are not the same file.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 150 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 6, 2026, and from our analysis. They are not legal advice.
Editorial analysis
160 questions, and the ground truth is a competition sheet
The corpus has a source that most benchmarks do not have. Questions come from the annual Global Fortune Teller Competition, for 2022 through 2025, and the raw sheets live under `data/raw/`. The normalized set in `data/data.json` holds 160 multiple-choice questions spread across twelve life aspects, with career, health, marriage, children and wealth named as examples and the rest elided.
Scoring is exact match against a ground-truth answer, not a judge model and not a rubric. That choice is what makes the numbers comparable across models without a second model in the loop, and it is also what makes the questions themselves the hard part: the answer is a fixed string that a practitioner published, and a model has to select it exactly.
Two file paths carry the whole dataset story. `data/data.json` is the question set, and `--year` narrows it to 2022, 2023, 2024 or 2025. The `--stats` flag prints dataset statistics, optionally filtered by the same year, so you can see what you are about to run before spending API calls on it. `--sample N` limits the run to the first N questions as a smoke test.
--astro injects a chart so the model is not graded on date arithmetic
The second file, `data/fortune_api_results.json`, holds pre-computed Bazi and Ziwei charts. They were produced with iztro and are keyed by `case_id`, which is what lets them be joined to `data.json` at runtime when `--astro` is set.
The stated purpose is isolation. With the flag on, the chart is injected into the prompt and the model does not have to derive it from the birth date, so a score reflects reasoning over the chart rather than the accuracy of date-to-chart conversion. That is a real methodological choice rather than a convenience: deriving a Bazi chart from a birth date is a deterministic computation with a well-known correct answer, and a model failing at it tells you about its arithmetic, not about fortune telling.
The flag is off by default, and the quick start calls `--cot` and `--astro` the recommended defaults, to be passed always and dropped only when you specifically want to ablate either effect. A plain run therefore measures the path the project does not recommend. Both flags are switches with no value: `--cot` prepends a chain-of-thought instruction to the prompt, `--astro` injects the charts.
A run with a slash in the model name goes somewhere different
Routing is inferred from the model name unless you override it, and the inference has two branches. A name containing a slash is treated as an OpenRouter identifier. A name without one has its provider inferred from the prefix, with the recognised prefixes being `gpt-*`, `claude-*`, `gemini-*`, `deepseek-*` and `doubao-*`. So `--model gpt-4o` and `--model openai/gpt-4o` do not reach the same place.
Leaving `--platform` unset is the documented path for OpenRouter, and four example lines use it:
python -m mingli_bench.cli --model openai/gpt-4o --year 2025 --cot --astro --max-workers 8
python -m mingli_bench.cli --model anthropic/claude-sonnet-4-6 --year 2025 --cot --astro
python -m mingli_bench.cli --model google/gemini-2.5-pro --year 2025 --cot --astro
python -m mingli_bench.cli --model deepseek/deepseek-r1 --year 2025 --cot --astro`--platform` accepts `openai`, `openrouter`, `anthropic`, `google`, `deepseek` and `doubao`, and passing it explicitly overrides the prefix inference. `--max-workers` defaults to 5, with 8 to 16 suggested when rate limits allow and a lower value when throttling appears. Two inspection flags exit without spending anything: `--list-models` prints supported model names and `--stats` prints dataset statistics.
Doubao needs an endpoint id before a key will do anything
Doubao is the one provider on the list that needs two values instead of one. The configuration block states that both the key and the endpoint id are required, with `DOUBAO_BASE_URL` pointing at the Volcengine Ark host and `DOUBAO_ENDPOINT_ID` taking an `ep-` value.
That second value has a consequence for routing. A versioned endpoint id such as `doubao-seed-2-0-pro-260215` does not match the prefix rules, which is the stated reason for passing `--platform` by hand, and the same example is also shown routed through an OpenAI-compatible gateway by setting the platform to `openai` instead:
# Native Doubao / Volcengine
python -m mingli_bench.cli \
--platform doubao --model doubao-seed-2-0-pro-260215 \
--year 2025 --cot --astro --max-workers 8
# Any model name, routed through an OpenAI-compatible gateway (set via OPENAI_BASE_URL)
python -m mingli_bench.cli \
--platform openai --model doubao-seed-2-0-pro-260215 \
--year 2025 --cot --astro --max-workers 8The requirements file carries three client libraries, openai, anthropic and google-generativeai. DeepSeek and Doubao are not in that list, which is consistent with both being reachable through the OpenAI-compatible path rather than through a client of their own.
The dotenv block and the .env.example in the tree are different files
The configuration step is one copy command, `cp .env.example .env`, after which empty and placeholder values are skipped automatically. `.env` is git-ignored, and `--env-file` points the CLI at a different file when you want a second key set. The trouble is that the block of settings shown in the README and the file that command copies are not the same.
The tree's .env.example spells out a performance section the README never reaches, and its Doubao base URL is a placeholder:
# Doubao / ByteDance
DOUBAO_API_KEY=your_doubao_api_key
DOUBAO_BASE_URL=your_doubao_base_url
DOUBAO_ENDPOINT_ID=your_doubao_endpoint_id
# ============================================
# Performance Settings
# ============================================
TIMEOUT=60000
MAX_TOKENS=81920
TEMPERATURE=0.0A user who fills in the key and leaves the rest alone gets a base URL of `your_doubao_base_url`, which is a bare hostname with no scheme and no host, while the README shows the real Ark endpoint at `https://ark.cn-beijing.volces.com/api/v3`. The other providers differ the same way, with the tree carrying an Anthropic base URL and a Google base URL that the README leaves out. Temperature of 0.0, a 60000 millisecond timeout and 81920 max tokens are the only record of those three settings anywhere, and they sit in the file you are told to copy rather than in the documentation. The README's own dotenv block breaks off inside a comment line before it reaches any defaults at all.
Category filters take Chinese names from an English command line
The dataset is described in English, but the filter takes the original names. `--categories` accepts the twelve aspects as 事业、健康、外貌、婚姻、子女、学业、官非、家庭、性格、灾劫、财运、运势, and the example given is `--categories 事业 婚姻`, which selects career and marriage.
The data table above it renders the same set in English, naming career, health, marriage, children and wealth and then trailing off with an ellipsis. So seven of the twelve categories have no English label in the documentation, and the one filter that lets you control what the model is actually asked about is driven by strings a reader cannot get from the prose.
`--shuffle-options` sits next to it and is worth pairing with any comparison run. It randomises the option order per question, which the documentation gives as a guard against position bias, a real risk when a model can do better by learning which letter is usually right.
Every run writes three artifacts under `--output-dir`, which defaults to `logs`: a results JSON with per-question predictions, scoring and aggregates, a summary text file with the headline numbers, and a directory of raw responses, one file per question. `--no-save` prints to the terminal and writes nothing.
Editorial conclusion
MingLi-Bench is a measuring stick, not a model, and it is worth running if your interest is how language models handle Chinese technical prose under exact-match scoring. Two things decide whether the number means anything: whether you pass --cot and --astro, because the recommended configuration is not the default one, and whether the keys in your .env point where you think they do, because the dotenv block in the README and the .env.example in the tree are not the same file. Do not read a score from a default run as a reasoning score, and treat the twelve category names as Chinese input even when the command line is English. The repository has no GitHub releases, its last commit is dated 2026-05-09, and the contact address given is [email protected].
Frequently asked questions
How many questions are in MingLi-Bench and where do they come from?
160 normalized multiple-choice questions across twelve life aspects, drawn from the annual Global Fortune Teller Competition for 2022 through 2025, with the raw sheets under data/raw/. Scoring is exact match against the ground-truth answer, and --year narrows a run to one of those four years.
What does the --astro flag do in MingLi-Bench?
It injects pre-computed Bazi and Ziwei charts from data/fortune_api_results.json into the prompt. Those charts were produced with iztro and are keyed by case_id so they can be joined to data.json at runtime, which means the model does not have to derive a chart from the birth date and the score reflects reasoning rather than chart conversion accuracy.
How do I configure API keys for MingLi-Bench?
Copy .env.example to .env with cp .env.example .env and fill in only the providers you plan to use, since empty and placeholder values are skipped automatically. The file is git-ignored, and --env-file points the CLI at a different one. Doubao needs both a key and an endpoint id, while OpenRouter takes one key.
Which providers can MingLi-Bench call?
The --platform flag accepts openai, openrouter, anthropic, google, deepseek and doubao. Without it, a model name containing a slash is treated as an OpenRouter identifier, and a name without one has its provider inferred from a prefix such as gpt-, claude-, gemini-, deepseek- or doubao-.
Where does MingLi-Bench write its results?
Three artifacts go under --output-dir, which defaults to logs: a results JSON holding per-question predictions, scoring and aggregates, a summary text file with the headline numbers, and a directory of raw model responses with one file per question. The --no-save flag prints to the terminal and writes no files.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/destinylinker-mingli-bench)