# WeiboAI/VibeThinker: a 1.5B and 3B dense reasoning model family for verifiable reasoning

> VibeThinker is a small dense reasoning model family from WeiboAI, post-trained with the Spectrum-to-Signal Principle and released under MIT. It targets math, competitive programming and STEM tasks where answers can be checked, and it does not ship an inference stack of its own.

**WeiboAI/VibeThinker** — Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B

- Repository: https://github.com/WeiboAI/VibeThinker
- Stars: 1,580 · Forks: 117
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/weiboai-vibethinker

## The problem VibeThinker takes on: reasoning quality without frontier-scale parameters

The project's stated premise is that small models are assumed to lack reasoning ability, and that this assumption is worth attacking directly. VibeThinker-1.5B is a 1.5B-parameter dense model, and VibeThinker-3B is a 3B-parameter dense model built on Qwen2.5-Coder-3B. Both are aimed at tasks with a reliable verification signal: mathematical reasoning, competitive programming, STEM reasoning, and instruction following with explicit constraints. That phrase, reliable verification signal, is the load-bearing part of the pitch. The models are not positioned as general chat assistants. They are positioned for problems where a correct answer exists and can be scored. The audience is therefore narrow and specific: researchers and engineers who want to run reasoning experiments on modest hardware, and teams that need a small model to sit inside a pipeline where the output is checked rather than trusted. The README reports that VibeThinker-1.5B surpasses the initial DeepSeek R1 model on AIME24 (80.3 vs. 79.8), AIME25 (74.4 vs. 70.0) and HMMT25 (50.4 vs. 41.7), despite R1 being over 400 times larger. Those are the project's own reported figures, not independent measurements.

## How the Spectrum-to-Signal Principle shapes training

The mechanism described in the README is a post-training pipeline, not an architectural change. The base model is dense and standard; the work happens after pre-training. For VibeThinker-1.5B the pipeline has two named stages. The SFT phase uses Two-Stage Diversity-Exploring Distillation, which the README says generates a broad spectrum of solutions. The RL phase then applies MaxEnt-Guided Policy Optimization (MGPO), described as amplifying the correct signal out of that spectrum. The name Spectrum-to-Signal follows from that split: first widen the space of candidate solutions, then concentrate on the ones that verify. VibeThinker-3B keeps the same principle but upgrades the surrounding machinery. According to the README, its pipeline combines curriculum-based supervised fine-tuning, multi-domain reinforcement learning, offline self-distillation, and instruction-oriented reinforcement learning, with strengthened data synthesis, quality filtering and curriculum learning in SFT. The 3B report also introduces Claim-Level Reliability Assessment (CLR), a test-time scaling strategy for answer-verifiable reasoning. CLR is inference-side, not training-side, and the README reports it raising AIME26 from 94.3 to 97.1 and HMMT25 from 89.3 to 95.4. The design bet is consistent across both sizes: spend the effort where verification exists, and let the verifier drive the optimization.

## Getting the VibeThinker weights: what the repository actually gives you

There is no pip install for the model itself and no bundled server. The README points to weight downloads on Hugging Face and ModelScope for both sizes, and the repository ships an eval/ directory alongside two PDF technical reports. The first real step is therefore fetching weights from one of those two hosting platforms. The README does not document a CLI, a port, an environment variable or a serving command, so any inference path you build is yours to choose. What the README does give is the exact model identifier for each size. For VibeThinker-1.5B it lists WeiboAI/VibeThinker-1.5B on Hugging Face and WeiboAI/VibeThinker-1.5B on ModelScope; for VibeThinker-3B it lists WeiboAI/VibeThinker-3B on both. Those identifiers are the only concrete handles the README provides, and they are what you pass to whatever loader you already use.

The README does not publish a recommended generation configuration, so sampling parameters, context length and prompt format are unspecified there. That is a real gap for a first run: a reasoning model's benchmark numbers depend heavily on how it is prompted and how many tokens it is allowed to generate, and the repository text does not pin those down. The eval/ directory is the place to look for the harness the authors used. Beyond that, treat the two PDFs in the repository root, VibeThinker-1.5B.pdf and VibeThinker-3B.pdf, as the primary specification, since the README's benchmark claims are drawn from the technical reports they correspond to.

## Where VibeThinker is the wrong tool

The clearest limitation is scope. Every capability claim in the README is attached to a verifiable benchmark: AIME24, AIME25, HMMT25, AIME26, BruMO25, LiveCodeBench v6, LeetCode contests. There is no reported result for open-ended generation, long-form writing, retrieval-augmented question answering, or agentic tool use. If your task has no automatic checker, the training signal this project is built around does not apply, and the benchmark numbers tell you nothing about whether the model will be good at it. The second limitation is packaging. The README lists downloads and a repository tree with eval/, figures/ and the PDFs. It does not document a quantized release format, a serving binary, an OpenAI-compatible endpoint, or an integration with any inference runner. Anyone expecting to pull a single artifact and start a server will have to build that layer. The third limitation is verification. The headline comparisons, including the DeepSeek R1 comparison and the $7,800 post-training cost figure, come from the project's own technical reports. They have not been reproduced here, and the README does not describe an independent evaluation. Treat them as claims to check against the eval/ harness rather than as settled facts. Finally, the README does not document rollback, version pinning between the 1.5B and 3B releases, or a changelog, so tracking behaviour across releases is on you.

## How VibeThinker differs from a general-purpose small model

The obvious alternative is to take a general instruction-tuned small model of similar size and use it as-is. The difference in approach is what the training optimizes for. A general instruct model is tuned across a wide distribution of tasks, with helpfulness and conversational quality as part of the objective. VibeThinker's post-training, as described, is organized around producing diverse candidate solutions and then reinforcing the ones that verify. That is a narrower objective, and it shows in what the project chooses to report: no chat benchmarks, no safety evaluations, no multilingual results. The trade is deliberate. You give up breadth for density of reasoning ability at a fixed parameter count. The other alternative is simply to use a much larger model and accept the cost. The README frames the comparison in those terms throughout, contrasting 1.5B and 3B against DeepSeek R1 at 671B and Kimi K2 at 1000B+. If your deployment has the memory and latency budget for a large model, the small-model argument weakens considerably, because the reported parity is on specific benchmarks and not across the board. The case for VibeThinker is strongest where the hardware ceiling is real and the task is checkable.

## Licence, maintenance and upgrade cost

The repository is MIT licensed, which is permissive and places few restrictions on reuse, modification or redistribution. The MIT text in the repository root governs the code and repository contents. The model weights are distributed separately through Hugging Face and ModelScope, and the README does not restate the licence terms for the weights in the text shown; check the model card on the hosting platform before you assume the repository licence carries over. This is a description of what the material says, not legal advice. On maintenance, the last push to the default branch was on 2026-08-14. The repository is not archived. The README's news section records a steady cadence: VibeThinker-1.5B open-sourced on 2025-11-11, VibeThinker-3B released on 2026-06-16, and the CLR paper released on 2026-08-12 with a separate code repository. There are no retrieved releases on the repository itself, so versioning appears to be handled through commits and through the model hosting platforms rather than through tagged releases. The upgrade cost is therefore mostly about weights: moving between 1.5B and 3B changes your memory footprint and your inference latency, and the README offers no guidance on migration between them. Budget for re-running your own evaluations after any weight update, because the repository does not publish a changelog of behavioural differences.

## Conclusion

Adopt VibeThinker if you need a small dense reasoning model for math, competitive programming or STEM tasks with checkable answers, and you are willing to assemble your own inference and serving path around the Hugging Face or ModelScope weights. Do not adopt it if you need a packaged runtime, a documented tool-calling interface, or a model for open-ended writing and chat. Before committing, verify three things: that the repository's eval/ directory covers the benchmark you care about, which of the two model sizes fits your hardware, and whether your licence obligations under MIT match how you intend to redistribute the weights.

## FAQ

### What is VibeThinker?

VibeThinker is a family of small dense reasoning models from WeiboAI, released under the MIT licence. The repository covers VibeThinker-1.5B, a 1.5B-parameter model, and VibeThinker-3B, a 3B-parameter model built on Qwen2.5-Coder-3B, both aimed at math, competitive programming and STEM reasoning where answers can be verified.

### What is VibeThinker 3B?

VibeThinker-3B is the 3-billion-parameter dense reasoning model in the family, built on Qwen2.5-Coder-3B and post-trained with an upgraded Spectrum-to-Signal pipeline that adds curriculum-based supervised fine-tuning, multi-domain reinforcement learning, offline self-distillation and instruction-oriented reinforcement learning. The README reports 94.3 on AIME26, 89.3 on HMMT25 and 80.2 Pass@1 on LiveCodeBench v6.

### How do I use VibeThinker?

Download the weights from Hugging Face or ModelScope using the identifiers WeiboAI/VibeThinker-1.5B or WeiboAI/VibeThinker-3B, then load them with your own inference code. The README does not document a CLI, a serving command or recommended generation settings, so the eval/ directory and the two PDF technical reports in the repository root are the reference points for reproducing the reported results.

## Sources

- [Issues](https://github.com/WeiboAI/VibeThinker/issues)
- [License: MIT](https://github.com/WeiboAI/VibeThinker/blob/main/LICENSE)
- [README](https://github.com/WeiboAI/VibeThinker/blob/main/README.md)
- [WeiboAI/VibeThinker on GitHub](https://github.com/WeiboAI/VibeThinker)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/weiboai-vibethinker
