DeepSeek-V3's 685B download, its RACE regression against V2, and a year-old repository
The repository metadata lists Python as its primary language. The metadata lists the MIT license. This article stays within the project description and details documented in the GitHub repository README.
At a glance
- What is it?
- DeepSeek-V3 is a Mixture-of-Experts model with 671B total parameters and 37B activated per token, published as a weights drop rather than a training pipeline. Its own benchmark table records it scoring below DeepSeek-V2 on RACE-Middle and RACE-High, and the repository's last push is dated 2025-08-28.
- Who is it for?
- DeepSeek-V3 fits a team that has already decided it can host 685B of weights and cares about MMLU-Pro, HumanEval, or MBPP behaviour. It does not fit a team choosing a checkpoint for multiple-choice reading comprehension, because the project's own table puts V3 below DeepSeek-V2 on both RACE splits.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Probably not. The repository last received commits 13 months ago, on August 28, 2025.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The 671B model is a 685B download, and 14B of that is the MTP module
The downloads table lists two artifacts, DeepSeek-V3-Base and DeepSeek-V3, both at 671B total parameters, 37B activated, and a 128K context length. A note under the table then corrects the arithmetic a reader would do next: the total size of the models on Hugging Face is 685B, which is 671B of Main Model weights plus 14B of Multi-Token Prediction module weights.
So the two numbers in circulation refer to different things. 671B is the main model. 685B is what you actually pull. The project does not present a flag, a config key, or a documented download variant that omits the MTP module, so a reader who does not want those 14B has no stated way to leave them on disk.
The MTP objective is described as beneficial to model performance and as usable for speculative decoding for inference acceleration, which is the reason to keep it. The catch sits in the same paragraph: MTP support is described as under active development within the community, with contributions and feedback welcomed. The acceleration path therefore depends on code this repository does not contain. DeepSeek-V3 cannot give you a faster inference loop out of the box, and the speedup that justifies 14B of extra weights is a thing the community has to finish.
Load balancing with no auxiliary loss, and nothing on the loss curve to read
The architecture section names two mechanisms above DeepSeek-V2. The first is an auxiliary-loss-free strategy for load balancing, which the project credits with minimizing the performance degradation that arises from encouraging load balancing. The second is the Multi-Token Prediction training objective. The framing of the first one is a claim about cost: the usual way to keep a Mixture-of-Experts model from collapsing onto a few experts is an auxiliary loss term, and that term competes with the objective the model is actually trained for. Removing it is presented here as a win rather than a simplification.
The cost of removing it is that the balancing signal leaves the loss curve. There is no term whose value tells you how evenly the router is distributing work, so expert collapse does not announce itself as a rising number in a training log. A reader watching a fine-tune has one fewer signal to watch and one fewer way to notice that a subset of experts is absorbing the traffic.
And for anyone reproducing the numbers, the repository does not state what replaces the auxiliary loss. If your implementation adds the conventional balancing term back, the model you train is not the model the evaluation table describes, and a gap between your scores and the published ones has an explanation before it has a bug.
2.788M H800 GPU hours, of which pre-training alone is 2.664M
The compute figures are given three ways and they are worth separating. The introduction states that DeepSeek-V3 requires only 2.788M H800 GPU hours for its full training. The pre-training section gives 2.664M H800 GPU hours for the 14.8T token pre-training run. The same section states that the stages after pre-training require only 0.1M GPU hours.
That arithmetic has a consequence for anyone budgeting a project of this shape. The expensive part is the base model, and the cheap part is everything that makes it usable, since the supervised fine-tuning and reinforcement learning stages together come to roughly a thirtieth of the pre-training bill. Teams that budget for the post-training work and underestimate the base run are off by an order of magnitude in the wrong direction.
Two supporting claims are made in the same voice. The training is described as remarkably stable, with no irrecoverable loss spikes and no rollbacks across the entire run. And the pre-training is FP8, through a mixed precision framework the project says it validated at this scale for the first time, with co-design across algorithms, frameworks, and hardware to nearly overlap computation and communication in cross-node MoE training. Those are claims about a specific cluster, and a reader without H800 hardware is not reading about a setup they can reproduce.
DeepSeek-V3 scores below DeepSeek-V2 on both RACE splits
The base-model table reports eighteen benchmark rows, and DeepSeek-V3 does not lead all of them. It takes ten, ties one, and loses seven. Two of the losses are the same task measured twice, and they are against its own predecessor.
On RACE-Middle, DeepSeek-V3 scores 67.1 against 73.1 for DeepSeek-V2 and 68.1 for Qwen2.5 72B. On RACE-High, it scores 51.3 against 52.6 for DeepSeek-V2 and 56.8 for LLaMA3.1 405B. The same model that leads MMLU-Pro at 64.4, HumanEval at 65.2, and MBPP at 75.4 sits at the bottom of a multiple-choice reading comprehension benchmark with answer candidates.
The remaining five losses are close calls on other rows: Pile-test BPB of 0.548 against 0.542 for LLaMA3.1 405B, HellaSwag of 88.9 against 89.2, PIQA of 84.7 against 85.9, WinoGrande of 84.9 against 86.3, and NaturalQuestions of 40.0 against 41.5. Several of those go to DeepSeek-V2 as well.
For a reader choosing a checkpoint, this is a selection criterion rather than a footnote. On this project's own numbers DeepSeek-V3 cannot beat DeepSeek-V2 at RACE, and a workload built around long passages with answer choices is exactly where a 671B model is supposed to pay off. Benchmark both against your own data before assuming the newer checkpoint carries the older one forward on every axis.
The tree is a weights drop with two license files and no training pipeline
The repository's top level holds eight entries: .github/, .gitignore, LICENSE-CODE, LICENSE-MODEL, README.md, README_WEIGHTS.md, figures/, and inference/. There is no data directory, no training script at the root, and no evaluation harness directory. The project's primary language is Python, and inference/ is the only place a program would live.
The two license files are the part that deserves a second look. The repository is recorded as MIT, and there is a LICENSE-CODE file that matches that description. There is a separate LICENSE-MODEL file beside it. A repository-level license field is a statement about the repository, and a reader who takes MIT as the answer for a 685B download from a model host has inferred something the tree does not say. The weights carry their own terms in their own file, and that is the file to read before storage, redistribution, or serving.
The consequence for anyone trying to learn from this project is a limit worth stating plainly. You can read the architecture claims, the parameter counts, the compute figures, and the evaluation table, and you can read the weight layout notes in README_WEIGHTS.md. You cannot read the code that produced any of those numbers, so the evaluation results are something you accept or go measure yourself rather than something you can rerun.
37B of 671B active per token, on an architecture V2 already used
DeepSeek-V3 is described as a Mixture-of-Experts model with 671B total parameters and 37B activated for each token, so the routing leaves a small fraction of the weights in play per step. The efficiency comes from the architecture it builds on: Multi-head Latent Attention and the DeepSeekMoE design, both of which the project says were thoroughly validated in DeepSeek-V2. The auxiliary-loss-free balancing strategy and the Multi-Token Prediction objective are the new parts.
The evaluation table keeps DeepSeek-V2 in the comparison columns for exactly this reason. DeepSeek-V2 is 236B total with 21B activated, listed as MoE, against Qwen2.5 72B and LLaMA3.1 405B, both listed as Dense with all 72B and all 405B activated. A reader comparing the three is comparing an MoE that touches 37B per token against dense models that touch everything.
What that table cannot tell you is how the 671B behaves under a serving configuration, and this repository does not report one. There is no throughput number, no latency figure, and no memory footprint anywhere in the README, only parameter counts and benchmark accuracy. The dense models in the comparison run 72B and 405B of arithmetic per token while V3 runs 37B, which is the argument for the design, but an argument is not a serving benchmark. Compute the memory yourself before planning hardware.
One release, a last push dated 2025-08-28, and MTP work pointed elsewhere
The repository has a single tagged release, v1.0.0, dated 2025-06-27. The last push is dated 2025-08-28. As of today that commit is more than a year old, and this repository is not where work on DeepSeek-V3 is visibly landing.
The project says as much without saying it. In the downloads section, MTP support is described as under active development within the community, with contributions and feedback welcomed. The post-training section attributes the reasoning behaviour to distillation from one of the DeepSeek R1 series models, which points at a separate line of work. Both signals describe improvements living outside the tree that holds the weights.
For a reader this changes what the repository is for. It is a fixed snapshot: a paper link, a parameter summary, a compute account, a benchmark table, and two download links on Hugging Face. DeepSeek-V3 cannot be expected to gain a new checkpoint, a new evaluation, or a new local-run recipe here, because nothing in its recent history points that way. If you need a maintained runtime, a current inference implementation, or a supported speculative decoding path, none of them are in this repository, and the community MTP work is the place the project itself redirects you.
Editorial conclusion
DeepSeek-V3 fits a team that has already decided it can host 685B of weights and cares about MMLU-Pro, HumanEval, or MBPP behaviour. It does not fit a team choosing a checkpoint for multiple-choice reading comprehension, because the project's own table puts V3 below DeepSeek-V2 on both RACE splits. Before committing, read LICENSE-MODEL separately from LICENSE-CODE, decide whether the 14B MTP module is worth carrying when its support lives outside this repository, and benchmark V2 beside V3 on your own task rather than on the aggregate.
Frequently asked questions
How big is DeepSeek-V3?
It has 671B total parameters with 37B activated for each token, and both DeepSeek-V3 and DeepSeek-V3-Base are published with a 128K context length. The download is larger than the parameter count suggests.
When was DeepSeek-V3 released?
The repository's only tagged release is v1.0.0, dated 2025-06-27. The last push to the repository is dated 2025-08-28, so the tree has not been updated in over a year.
What are the parameters of DeepSeek-V3?
671B total with 37B activated per token, using Multi-head Latent Attention and the DeepSeekMoE architecture, both validated in DeepSeek-V2. The Hugging Face download totals 685B because it adds 14B of Multi-Token Prediction module weights to the 671B main model.
how to install deepseek v3 locally
Both DeepSeek-V3 and DeepSeek-V3-Base download from Hugging Face, and the project's table of contents has a dedicated How to Run Locally section that it refers readers to for step-by-step guidance. The repository's top level also carries an inference/ directory, and README_WEIGHTS.md covers the main model weights and the MTP modules.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/deepseek-ai-deepseek-v3)