# pyKT tags a 1.0.0 release but setup.py still declares 0.0.38

> A PyTorch benchmarking library for deep learning knowledge tracing models, wrapping more than thirty published approaches over seven or more datasets and five prediction scenarios. It is well cited and thoroughly listed, and its packaging metadata has drifted away from its own tags in a way that changes what you actually install.

**pykt-team/pykt-toolkit** — pyKT: A Python Library to Benchmark Deep Learning based Knowledge Tracing Models

- Repository: https://github.com/pykt-team/pykt-toolkit
- Website: https://pykt.org
- Stars: 441 · Forks: 137
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/pykt-team-pykt-toolkit

## The newest tag is 1.0.0, but the installable package is 0.0.38

There are three recent releases and they tell a slightly awkward story.

v0.0.37 was published in July 2022, v0.0.38 in October 2022, and v1.0.0 in February 2023. The 1.0.0 release is not titled with a version note; its release name is merge models.

The package metadata has not followed. setup.py declares:

```python
    name="pykt-toolkit",
    version="0.0.38",
```

Since pip builds from that declaration, installing the documented way gets you 0.0.38, not the 1.0.0 tree. The default branch has moved well past both.

So there are three versions in play for anyone starting today: what they get from the index, what the 1.0.0 tag contains, and what is on main. The 0.0.38 line and the 1.0.0 line are separated by four months of work whose release name suggests the change was structural rather than incremental.

The repository is not archived and the last push is dated 2026-09-22, so none of this is a dead project. It is a project whose packaging metadata stopped tracking its own history.

## The README is read into long_description and then immediately discarded

The top of setup.py reads the README and assigns it to a variable:

```python
with open("README.md", "r", encoding="utf-8") as fh:
    long_description = fh.read()
```

Then the setup call does not use it. The long_description keyword receives a single-line string, the project's short description, while long_description_content_type is declared as text/markdown.

The result is that the PyPI page renders one sentence of prose with a markdown content type declared, and the README content that was just read from disk never reaches the index.

This is a small thing on its own. It matters more as evidence, because it is the kind of drift that a project accumulates when the release process runs on a schedule rather than on a change. The same drift explains the version mismatch, and both are visible from the repository without running anything.

The rest of the metadata is thin too. Three classifiers are declared, for Python 3, for MIT and for operating system independence, and there is no Development Status classifier, so the index carries no signal about maturity.

## The documented environment is Python 3.7.5 with a legacy activation command

Installation is two blocks, and the first one pins an interpreter that has been end-of-life for years.

```bash
conda create --name=pykt python=3.7.5
source activate pykt
```

The activation syntax is the older conda form. The current command is conda activate without source, and using the old form against a recent conda installation is a common source of shell errors.

The package metadata is looser than the documentation. python_requires is set to 3.5 or newer, so anything from 3.5 upward is declared acceptable, while the tutorial asks for exactly 3.7.5.

Nothing in the repository says which combinations are actually tested, and the space between the two numbers is wide. A dependency floor of torch 1.7.0 sits under it, and modern PyTorch does not install on 3.7.

So the practical reading is that the tutorial's 3.7.5 is the tested configuration for the era of the library, and the 3.5 floor in the metadata is a declaration rather than a claim. Anyone on a current interpreter should expect to resolve the dependency set themselves.

## wandb is a required dependency, so logging is not opt-in

The install_requires list has six entries and one of them is a hosted experiment-tracking service:

```python
    install_requires=['numpy>=1.17.2','pandas>=1.1.5','scikit-learn','torch>=1.7.0','wandb>=0.12.9','entmax'],
```

Weights and Biases is therefore a hard requirement of the library, not an extra you add when you want telemetry. Anyone installing pyKT is installing a client for a third-party service.

That choice runs through the whole examples directory, where training scripts are named for the tracker rather than for the model: wandb_akt_train.py, wandb_atdkt_train.py, wandb_atkt_train.py, wandb_cgmkt_train.py, wandb_cskt_train.py, wandb_datakt_train.py, wandb_deep_irt_train.py and wandb_denoisekt_train.py, among others.

entmax is the least obvious entry. It provides the sparse attention transformation used by several of the models in the catalogue, which is why a benchmark library pulls an attention primitive as a core dependency rather than leaving it to each model.

There is no optional-dependency group in the visible metadata, so there is no documented way to install the models without the tracker.

## One training script per model, wrapped in scripts that generate and merge runs

The examples directory is where the design of the benchmark becomes visible, and it is organised around two ideas.

The first is one script per model, named wandb_<model>_train.py, so that every entry in the catalogue is launched the same way and can be compared on the same harness. The second is that the runs are then swept. Alongside the per-model scripts sit generate_wandb.py, generate_ab_wandb.py for paired comparisons, generate_splitpred.py, merge_wandb_results.py, and check_wandb_status.ipynb as a notebook for inspecting what has been submitted.

There is also a seedwandb/ directory and a competitions/ directory, which suggests the sweeps and the challenge runs are kept separate.

Shell drivers sit on top of the Python. run_all.sh, multi_run_all.sh, dataprocess.sh and extract_raw.sh each stage one part of the pipeline: preprocessing, extraction of raw results, and launching everything.

Taken together this is a fairly complete answer to the question of how thirty models get compared fairly. The cost is that the surface area is large and almost all of it assumes the tracker is present.

## Training scripts exist for models the thirty-two item reference list does not name

The README credits eighteen upstream project repositories and lists thirty-two papers, which is the transparency the project is built on.

The list is not perfectly maintained. One entry, FlucKT on cognitive fluctuations in attention networks, is present but commented out, so it is neither endorsed nor removed.

More interestingly, the training scripts do not map one-to-one onto the paper list. wandb_cgmkt_train.py, wandb_deep_irt_train.py and wandb_datakt_train.py name models that do not correspond to any of the thirty-two titles given, and the paper list has been stable across the three releases.

So the implemented surface is wider than the documented catalogue. That is not necessarily a problem, since a benchmark that runs an extra model is more useful than one that refuses to, but it does mean the reference list is a lower bound on what the code can do rather than a complete index.

The credit list itself is thorough in a way worth acknowledging: DKT through DenoiseKT, from the foundational models to the recent ones on length generalisation and question-centric representations, each paired with the public repository it came from.

## Tuning results are a Google Drive link, not files anyone can diff

Reproducibility here rests on one external asset.

The hyperparameter tuning results for all the models on all the datasets are published as a Google Drive folder linked from the README. Nothing in the repository describes the search space, the number of trials, or the selection rule, and the folder is not versioned with the code.

That is the single most fragile dependency in an otherwise self-contained project. A model configuration lives in configs/, preprocessing code lives in data/ and pykt/, tests live in tests/, and documentation lives in docs/, with build.sh at the root. The one thing that cannot be reconstructed from the checkout is the set of hyperparameters the reported numbers came from.

The scholarly anchor is a paper rather than a documentation site. The citation is an arXiv preprint, 2206.11460, published in the Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, authored by Liu Zitao, Liu Qiongqiong, Chen Jiahao, Huang Shuyan, Tang Jiliang and Luo Weiqi.

A Datasets and Benchmarks paper is the right venue for a library like this, and it is also an implicit statement that the artifact, not the model, is the contribution.

## Conclusion

This fits a research group reproducing or comparing knowledge tracing results who wants every baseline in one place with shared preprocessing, rather than reimplementing thirty papers. It does not fit a production service, since the dependency floor is old, the logging client is a hard requirement, and there is no stability classifier in the package metadata. Two things to check before you build on it. First, pip resolves to what setup.py declares, so you will get 0.0.38 even though a 1.0.0 tag exists, and that tag is titled with a merge rather than a release. Second, the hyperparameter tuning results are a Google Drive link rather than files in the repository, so plan for that link to be the weakest link in your reproduction. The last push is dated 2026-09-22.

## FAQ

### What is pyKT?

A Python library built on PyTorch for training and benchmarking deep learning knowledge tracing models. It provides standardised preprocessing across more than seven datasets, five prediction scenarios, and more than ten commonly compared approaches, with a longer catalogue of thirty-two cited papers.

### Which version of pyKT does pip install?

The one declared in setup.py, which is 0.0.38. A v1.0.0 tag exists from February 2023 and is titled merge models, but the package metadata still declares 0.0.38, so the index and the tag are different trees.

### How do I install pyKT?

Create a conda environment with python=3.7.5, activate it with source activate pykt, then run pip install -U pykt-toolkit -i https://pypi.python.org/simple. The package declares python_requires of 3.5 or newer and a torch floor of 1.7.0.

### Does pyKT require Weights and Biases?

Yes, wandb 0.12.9 or newer is listed in install_requires alongside numpy, pandas, scikit-learn, torch and entmax. There is no documented optional-dependency group, and the per-model training scripts are named wandb_<model>_train.py.

### Where can I find pyKT hyperparameter tuning results?

They are published as a Google Drive folder linked from the README, covering all models on all datasets. The results are not stored in the repository, so the search space and selection rule are not versioned alongside the code.

### How do I run all pyKT models for comparison?

The examples directory holds one training script per model named wandb_<model>_train.py, plus generate_wandb.py, generate_ab_wandb.py for paired runs and merge_wandb_results.py for combining them. Shell drivers including run_all.sh, multi_run_all.sh, dataprocess.sh and extract_raw.sh sit on top.

## Sources

- [License: MIT](https://github.com/pykt-team/pykt-toolkit/blob/main/LICENSE)
- [Project website](https://pykt.org)
- [pykt-team/pykt-toolkit on GitHub](https://github.com/pykt-team/pykt-toolkit)
- [README](https://github.com/pykt-team/pykt-toolkit/blob/main/README.md)
- [Releases](https://github.com/pykt-team/pykt-toolkit/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/pykt-team-pykt-toolkit
