MiniMind's 64M model, and what its two-hour claim covers
đź§ Train a 64M-parameter LLM from scratch in just 2h!
At a glance
- What is it?
- MiniMind is a Python repository that trains a model of roughly 64M parameters from scratch in native PyTorch, covering pretraining, SFT, LoRA, RLHF-DPO, RLAIF and agentic RL. It reads as a teaching codebase first: every algorithm is written by hand, and the headline two-hour number covers one SFT epoch on a single 3090.
- Who is it for?
- Adopt MiniMind if you want to read a complete LLM training pipeline line by line and are willing to install PyTorch yourself, since requirements.txt leaves torch commented out and the two-hour figure covers one SFT epoch on a 3090 rather than pretraining. Do not adopt it to run a production model, and do not plan around a minimind-v1 checkpoint, because that series is offline and weights from before 2025-04-26 no longer load directly.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 9 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Two hours means one SFT epoch on a single 3090
Two hours is the number this project leads with, and it covers less than most readers assume. The README states that the two hours refer to the SFT stage running 1 epoch on a single NVIDIA 3090, and that the 3 yuan figure is the GPU rental cost for that same window. Pretraining is not inside that number, and neither are the DPO or RLAIF stages with their PPO, GRPO and CISPO algorithms, each of which needs its own generation and training run. What the figure describes is finishing one stage of the pipeline cheaply, not producing a chat model in an afternoon. The code is Apache 2.0 and free to use, the last push to the repository was on 2026-09-22, and two releases exist: v2 for the MiniMind docs and minimind-v1 on ModelScope.
requirements.txt ships with torch commented out
PyTorch is missing from the dependency list, and that is the first thing a new user trips over. requirements.txt pins its packages exactly, from datasets==3.6.0 through transformers==4.57.6 to trl==0.13.0, but the four lines that would make a trainer runnable are commented out:
# torch==2.6.0
# torchvision==0.21.0
# peft==0.7.1
# matplotlib==3.10.0So install PyTorch yourself, matched to your own CUDA build, before the rest. The layout tells you what you are looking at: dataset/ holds the data, model/ the Dense and MoE definitions, trainer/ the training code, scripts/ the inference and conversion helpers, eval_llm.py the evaluation entry point. The same requirements file also pulls in Flask, fastapi, uvicorn, streamlit, wandb, swanlab, modelscope, jieba and sentence_transformers, which makes a first install much heavier than a model library needs to be.
The main line is Qwen3 shaped, at 64M and 198M-A64M
The current architecture follows Qwen3 rather than inventing its own shape. The 2026-04-01 release published minimind-3 at about 64M parameters and minimind-3-moe at 198M-A64M, rewrote the structure, tokenizer, training chain, inference interface and default configuration together, aligned the main line with the Qwen3 and Qwen3-MoE ecosystem, and removed the shared expert design. Earlier weights show how fast that line moves: minimind2 at 104M, minimind2-small at 26M and minimind2-moe at 145M all date from 2025-04-26, while the v1 family at 108M, 26M and 4Ă—26M dates from 2024. Default training data moved to pretrain_t2t(_mini).jsonl, sft_t2t(_mini).jsonl, rlaif.jsonl, agent_rl.jsonl and agent_rl_math.jsonl, and standalone train_reason.py was deleted, with thinking now handled by the chat_template and an open_thinking switch.
Tool use moved into the SFT data, so your own data loses it
Tool calling is no longer a separate recipe you add at the end. In the 2026-04-01 entry, toolcall capability was mixed into the sft_t2t and sft_t2t_mini main line data, so the default full_sft run already has basic Tool Call ability, and the standalone reasoning trainer disappeared with it. The bill arrives the moment you swap in your own corpus: the calls are template markers, <tool_call> and <tool_response> alongside <think>, so data without them yields a model that never emits them no matter how the template is written. Agentic RL keeps its own script, train_agent.py, running GRPO or CISPO across multi-turn tool use, and the RLAIF and Agentic RL pipelines were refactored so the rollout engine is decoupled and the generation backend can be replaced. scripts/chat_api.py is the new inference example.
Everything is native PyTorch, and that is both the pitch and the bill
Every core algorithm is implemented from zero rather than borrowed from a framework, and the README argues the case explicitly against libraries such as transformers, trl and peft, which expose high level interfaces where a dozen lines can load a model, load a dataset, run inference and run reinforcement learning, at the cost of hiding the implementation. MiniMind still stays compatible with those libraries, adds llama.cpp, vllm and ollama on the inference side and Llama-Factory on the training side, so an escape route exists when the hand written code stops being an advantage. The cost shows up when something breaks. With no abstraction layer between you and the hardware, the training loop, the KV cache and the DPO objective are yours to debug, which is why the 2025-10-24 entry lists normalised code and fixed known bugs as a headline item.
Pre-2025-04-26 weights no longer load, and Llama position encoding is why
Checkpoint compatibility is the sharpest edge in this repository, and the reason is stated rather than buried. To stay compatible with llama.cpp and vllm, that update stopped supporting direct loading of older models, because Llama position encoding differs from MiniMind's and the mapped QK values come out different; the minimind2 series recovered those checkpoints through weight mapping plus a fine-tuned QKVO linear layer calibration. The same entry abandons maintenance of the entire minimind-v1 series and takes it offline in the repository. If you pinned minimind-v1 at 108M from 2024-09-01, or v1-small and v1-moe, you are on a retired line, and the only conversion helper the changelog documents, scripts/convert_model.py, exists to merge LoRA weights into a complete model rather than to revive an old checkpoint.
Datasets went from bin to CSV to jsonl, and old tutorials rot
The data pipeline changed three times, and the file names are the record. In 2024-09-27 the project gave up preprocessing pretrain data into .bin files, accepting a small loss of training speed to keep the text intact, and the output became pretrain_data.csv. By 2025-02-09 the preprocessing step was removed altogether and the format standardised on jsonl, with the stated aim of ending the confusion around dataset downloads. The rlaif-mini.jsonl set added on 2025-10-24 holds 10,000 rows sampled at random from the SFT data, and the DPO set was simplified and given Chinese data. Anyone following a tutorial written before 2025 has to check which of these formats the code under dataset/ and trainer/ still expects, because a mismatched file name fails at load time rather than degrading quietly.
Vision, omni, diffusion and linear attention are not in this tree
Four of the project's directions sit outside the repository you would clone. Vision moved to the twin project MiniMind-V in 2024-10-05, the omni model is MiniMind-O, and a discrete diffusion language model plus a linear attention model are tracked as experimental extensions in the project's discussions, with the note that both can be further trained from the main autoregressive checkpoint. None of it appears in the tree, which holds dataset/, model/, trainer/, scripts/, eval_llm.py and the images used by the README. For evaluation you get eval_llm.py and support for C-Eval, C-MMLU and OpenBookQA, and for long context a YaRN based RoPE extrapolation; how to run either one is not described in the README, so budget reading time for the code under trainer/ before you plan a schedule.
Editorial conclusion
Adopt MiniMind if you want to read a complete LLM training pipeline line by line and are willing to install PyTorch yourself, since requirements.txt leaves torch commented out and the two-hour figure covers one SFT epoch on a 3090 rather than pretraining. Do not adopt it to run a production model, and do not plan around a minimind-v1 checkpoint, because that series is offline and weights from before 2025-04-26 no longer load directly. Before starting, confirm which release you are following and check that your own data carries the <tool_call> and <think> markers if you expect tool use.
Frequently asked questions
What is a MiniMind?
MiniMind is a set of open source model structures and training code for very small language models, the main line being a model of roughly 64M parameters trained from scratch rather than fine-tuned. The repository covers the whole pipeline: pretraining, SFT, LoRA, RLHF-DPO, RLAIF with PPO, GRPO and CISPO, tool use and agentic RL.
What type of AI model is MiniMind?
The main line is an autoregressive language model, published as a Dense version of about 64M parameters and a Mixture of Experts version at 198M-A64M, both aligned with the Qwen3 and Qwen3-MoE ecosystem. Vision, omni, discrete diffusion and linear attention variants are developed as separate projects.
Is MiniMind safe and secure?
The repository is Apache 2.0 licensed training code with a pinned requirements.txt, and no hosted service is described. The top level holds no security policy file, only CODE_OF_CONDUCT.md and LICENSE, so what you run is your own PyTorch installation and your own copy of the datasets.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/jingyaogong-minimind)