Open-source project
Calix-L/DanKS avatar
Calix-L/DanKS

DanKS, three generations of GuanDan agents sharing one rules engine

RL‑Empowered Small‑Scale Competitive Guandan Agent

638 stars15 forksPythonApache-2.0

At a glance

What is it?
DanKS keeps every generation of its GuanDan agent in one tree so you can read the progression from NumPy scoring to a PPO-trained policy. Here is how the packages divide, and what each refuses to do.
Who is it for?
DanKS fits someone studying how a card-game agent improves across generations rather than just using the best checkpoint, because the older code stays readable next to the newer code instead of being rewritten away. The shared engine plus separate per-generation environments is the arrangement that makes that possible.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A large action space becomes a bounded Top-K problem

The whole design is one sentence: turn a large, structured action space into a compact policy decision. The policy receives an encoded information state, which is the visible hand, the public action history, the legal actions, and seat-aware game context. Budgeted decomposition search then retrieves structured candidates and summarises each one by length, pairs, sequences, suits, gaps, and remaining-hand structure. A shared encoder combines the state, the candidate, and those structural features; an actor ranks the valid candidates and a critic estimates the state value. The fourth step is learning from self-play, where trajectories produce generalised advantage estimates for clipped PPO updates. The constraint stated on that last step is the important one: training improves the selector without expanding the inference-time candidate budget, so the work happens offline and the search stays bounded at play time.

Four generations in one tree, each with its own entry point file

The generations table names a main idea, what it adds, and a specific file to read first for each. V1 is structural retrieval with candidate scoring and a NumPy selector. V2 is learned selection, adding broader action generation and an ONNX selector. V3 is memory-aware policy learning, adding card memory, candidate coverage, recall, team belief, and PPO. V3Pro is inference-time refinement of V3, adding an asset gate, equivalent-play rules, and verified endgame search. Each of those points at a concrete module path rather than a directory, so the reading order is prescribed. V1, V2, and V3 ship as separate packages, and the instruction is to give each generation its own environment so that the shared package name, the feature schema, and the checkpoint format stay aligned across versions.

V1 scores candidates in NumPy, V2 exports a selector, V3 learns

The progression between the first two generations is about where the work happens. V1 does its selection with NumPy, which means no learned component at all and a selector you can read end to end. V2 keeps the structural retrieval idea but replaces hand-scored selection with a learned one, exports it as ONNX, and widens the action generation so there is more to select from. That export format is the detail that matters for anyone reproducing results: ONNX is a portable artefact, so a V2 selector runs outside Python. V3 is where the memory and learning arrive, and where the training pipeline lives. Each generation advances the features, the retrieval strategy, and the policy implementation inside the same four-step frame, which is what makes them comparable rather than three separate projects.

V3Pro is a source-only extension, not a fourth network

The newest row is scoped tightly on purpose, and the framing matters if you are deciding whether to install it. V3Pro is described as an optional source-only extension of V3, explicitly not a fourth network and not a replacement training pipeline. It installs alongside V3 under a different package name, and the original V3 model, its features, and its PPO implementation stay where they are. Its decision order is fixed and linear: retrieve candidates, apply the asset mask, run the frozen V3 policy, apply the equivalent-play rules, and finish with verified endgame refinement. Because the policy is frozen, the extension changes what gets chosen rather than what the network believes, which means it can be added or removed without retraining anything.

Each V3Pro component states what it refuses to do

The extension's documentation is unusually disciplined about limits, and the limits are the specification. The asset gate masks an ordinary candidate when a same-type, same-strength alternative spends fewer wildcards or breaks fewer natural bombs without worsening the guarded metrics. It is explicitly not a blanket wildcard ban: the original mask is preserved for immediate finishes and for opponents holding one or two cards, and matching singles or pairs are not banned outright. The equivalent-play rules keep the physical triple and swap in a smaller pair only when the residual partition is equivalent, conservatively choosing a smaller sufficient follow bomb, with no arbitrary triple replacement, no lead-bomb substitution, and no structural sacrifice. The endgame search enumerates the public hidden cards, continues with the frozen policy, applies exact minimax, handles ties only with a complete proof, and recovers candidates in isolation.

The endgame search stops at sixteen cards and 128 allocations

That last component has hard numeric bounds, and they are the ones to check against your position. It only applies when at most sixteen cards remain in total. Between eleven and sixteen cards it enumerates at most 128 hidden-card allocations. And when verification cannot be completed, it keeps the base action rather than substituting a partial result. Those three rules together mean the extension degrades to the plain V3 decision in exactly the cases where exact search would be too expensive, rather than degrading to a guess. The tie handling is the strictest part: a tie is only broken when there is a complete proof, so a position that looks like a tie but cannot be proved stays a tie. None of these bounds is described as adjustable, so they read as design decisions rather than defaults.

The root package is the engine, and it has no dependencies

The packaging tells you what is a library and what is an experiment. The root project is named for the engine, not for the agent, and its description says three generations of agent code and a shared Python game engine. Its declared package list is two modules, the top-level game package and its engine subpackage, and its dependency list is empty. So installing the root gives you the rules engine with nothing else, which is the correct shape for something the versions all sit on top of. The version packages are installed separately and directly from the tree, one editable install per generation. Development extras are a test runner and, for interpreters below 3.11, a TOML reader backport, and the test configuration is minimal: terse output plus one directory to look in.

The CPU quickstart pins one torch build from one wheel index

The shortest documented path runs the newest generation on CPU and assumes Python 3.11 with a POSIX shell:

bash
git clone https://github.com/Calix-L/DanKS.git
cd DanKS
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e versions/v3
python -m pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cpu
python examples/retrieval_quickstart.py --version v3
python examples/v3_model_smoke.py

The framework is installed from the vendor's CPU wheel index at an exact version rather than from the default index, which is what keeps the CPU build small and predictable. Note the small mismatch worth knowing: the package metadata declares support from 3.10 while the quickstart asks for 3.11 specifically, and the Windows activation command is a different path entirely. Accelerator setups, the older generations, native kernels, and development installs are all deferred to a separate installation reference rather than covered here.

Editorial conclusion

DanKS fits someone studying how a card-game agent improves across generations rather than just using the best checkpoint, because the older code stays readable next to the newer code instead of being rewritten away. The shared engine plus separate per-generation environments is the arrangement that makes that possible. It is a poor fit if you want one clean install of the strongest agent, because the root package installs the engine and nothing else, and it is a poor fit if you cannot accept a research codebase with five runnable examples rather than a maintained library. Before you start, read the installation reference for the accelerator path rather than assuming the CPU quickstart covers your hardware.

Frequently asked questions

What are the three generations of DanKS?

V1 is structural retrieval with a NumPy selector. V2 is learned selection with broader action generation and an ONNX selector. V3 is memory-aware policy learning with card memory, candidate coverage, recall, team belief, and PPO. A fourth row, V3Pro, refines V3's decisions at inference time without retraining.

Is V3Pro a separate model I have to train?

No. It is an optional source-only extension of V3 that installs alongside it under a different package name, and the frozen V3 policy stays where it is. Its decision order is retrieval, asset mask, frozen V3, equivalent-play rules, then verified endgame refinement, so it changes what gets chosen rather than what the network believes.

What does pip install from the DanKS repository root?

The shared game engine only. The root project declares two packages, the game package and its engine subpackage, and an empty dependency list. Each generation is then installed separately and directly from the tree with an editable install pointed at its version directory.

How do I run DanKS on a CPU?

Follow the quickstart, which assumes Python 3.11 and a POSIX shell. It clones, creates and activates a virtual environment, installs the V3 version package in editable mode, then installs torch at an exact version from the vendor's CPU wheel index. CUDA and Ascend NPU paths are documented separately in the installation reference.

When does the DanKS endgame search give up?

It applies only when at most sixteen cards remain in total, and between eleven and sixteen cards it enumerates at most 128 hidden-card allocations. When verification cannot be completed it keeps the base action, and a tie is only broken when there is a complete proof. None of those bounds is described as configurable.

Official sources

  1. Calix-L/DanKS on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/calix-l-danks.svg)](https://hysenlabs.com/projects/calix-l-danks)