Model or dataset
huggingface/smollm avatar
huggingface/smollm

SmolLM: Hugging Face's Family of Small, Fully Open Language and Vision Models

Everything about the SmolLM and SmolVLM family of models

3,913 stars317 forksPythonApache-2.0

At a glance

What is it?
The huggingface/smollm repository is the central resource for the SmolLM and SmolVLM model families from Hugging Face. SmolLM3, the current flagship, is a 3B parameter language model trained on 11 trillion tokens with multilingual support and dual-mode reasoning. SmolVLM extends the family to vision-language tasks.
Who is it for?
SmolLM3 is worth evaluating for teams that need a small, fully open model for inference on-device or on modest hardware, particularly if multilingual support across English, French, Spanish, German, Italian, and Portuguese is a requirement. The 3B model runs on consumer GPUs when loaded in standard precision.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 7 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the SmolLM Family Is and Who It Is For

The huggingface/smollm repository covers two related model families. SmolLM (now at version 3) is a series of language models designed to run efficiently on-device while maintaining competitive performance relative to other models in the same size range. SmolVLM extends the approach to vision-language tasks, enabling image understanding alongside text generation.

The repository targets three kinds of users. Developers who want to run a capable language model locally without the compute requirements of a 7B or larger model will find SmolLM3 a practical starting point. Researchers who need full reproducibility will benefit from the public training details: SmolLM3 is described in the README as fully open, with open weights, public data mixtures, and published training configurations. Organizations with multilingual requirements can use the model's built-in support for six languages without fine-tuning.

The repository structure groups resources by modality: a text/ directory covers SmolLM3, SmolLM2, and SmolLM1; a vision/ directory covers SmolVLM; and a tools/ directory holds shared inference utilities, local inference scripts, and the smol_tools package.

SmolLM3: What 3B Parameters at 11 Trillion Tokens Delivers

SmolLM3 is a 3B parameter model trained on 11 trillion tokens. The README states it outperforms Llama 3.2 3B and Qwen2.5 3B and remains competitive with larger 4B alternatives including Qwen3 and Gemma3. These comparisons are stated in the README; task-specific benchmark details are in the linked blog post rather than the repository itself.

The model includes several features that distinguish it from earlier SmolLM versions. Dual-mode reasoning supports both thinking and no-think modes, which the README indicates are available in the instruct variant. Long context support extends to 128k tokens through NoPE (No Position Encoding) and YaRN scaling. Multilingual support covers English, French, Spanish, German, Italian, and Portuguese without requiring language-specific fine-tuning.

SmolLM3 is based on the transformer architecture and is loaded through the standard Hugging Face Transformers library. The base model is available at HuggingFaceTB/SmolLM3-3B-Base and the instruction-tuned variant at HuggingFaceTB/SmolLM3-3B.

Loading and Running SmolLM3

SmolLM3 loads through the standard Transformers AutoModelForCausalLM interface:

python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "HuggingFaceTB/SmolLM3-3B"
device = "cuda"

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
).to(device)

prompt = "Give me a brief explanation of gravity in simple terms."
messages_think = [
    {"role": "user", "content": prompt}
]

text = tokenizer.apply_chat_template(
    messages_think,
    tokenize=False,
    add_generation_prompt=True,
)

The `apply_chat_template` call formats the input for the instruct model. Using `device = "cpu"` is an alternative for machines without a GPU, though inference will be slower.

For SmolVLM, the vision-language model uses a different processor class:

python
from transformers import AutoProcessor, AutoModelForVision2Seq

processor = AutoProcessor.from_pretrained("HuggingFaceTB/SmolVLM-Instruct")
model = AutoModelForVision2Seq.from_pretrained("HuggingFaceTB/SmolVLM-Instruct")

SmolVLM handles multiple images in a single conversation and performs tasks including visual question answering, image description, and visual storytelling according to the README.

SmolLM3 vs Comparable Small Models

The README names Llama 3.2 3B, Qwen2.5 3B, Qwen3, and Gemma3 as comparisons. The key differentiator the README emphasizes is transparency: the training dataset mixture and training configurations are public, which neither Llama 3.2 nor Qwen2.5 provide at the same level of detail.

Llama 3.2 3B from Meta AI is another small open model in the same size class. Meta releases model weights but does not publish the full pretraining data composition. The structural difference is that SmolLM3 is built on top of fully public datasets including the SmolLM3 pretraining collection, FineMath (a public mathematics pretraining dataset), and FineWeb-Edu (a public educational content dataset), making the training fully reproducible.

The SmolLM2 paper is available at arXiv:2502.02737v1 and covers the predecessor model's training methodology. The SmolLM3 blog post provides more recent training details. The instruct variant uses the SmolTalk dataset for instruction tuning, which is also published on Hugging Face.

For teams that prioritize multilingual coverage, SmolLM3's six-language support without fine-tuning is a practical advantage over models that require language-specific variants or fine-tuning for non-English tasks.

Limitations and What the Repository Does Not Cover

The repository provides model loading code and links to Hugging Face model pages but does not include quantization tooling, deployment server code, or inference optimization beyond the standard Transformers pipeline. Teams that need GGUF, AWQ, or GPTQ quantized versions must look for community quantization efforts on the Hugging Face hub.

SmolVLM is a compact model and the README notes it is designed for efficient on-device operation, but does not specify memory requirements for different image resolutions or batch sizes. Engineers integrating SmolVLM into applications should run their own benchmarks for their specific image dimensions.

The 3B size means SmolLM3 is less capable than 7B or larger models on complex reasoning tasks. The dual-mode reasoning support helps with structured thinking, but for tasks that require deep multi-step reasoning, larger models remain more capable. The README's position is that SmolLM3 is competitive at the 3B scale and with some 4B models, not that it competes with 7B or 70B class models.

The last push to the repository was on 2026-09-23, indicating active ongoing development.

Training Data, License, and Reproducibility

SmolLM3 is trained on a publicly documented data mixture. The pretraining datasets are published as a collection on Hugging Face under HuggingFaceTB/smollm3-pretraining-datasets. The instruction-tuning uses the SmolTalk dataset, also public. The mathematics pretraining uses FineMath. Educational content comes from FineWeb-Edu.

This level of transparency is the main argument the README makes for SmolLM3 as a research model. Teams that need to understand what data a model was trained on, either for reproducibility or for compliance reasons, can trace SmolLM3's training lineage through public sources.

The repository is licensed under Apache-2.0, which permits commercial use, modification, and redistribution. The model weights on Hugging Face follow their own license terms; verify those separately before commercial deployment.

The smol_tools package under tools/smol_tools/ provides lightweight AI-powered utility scripts built on top of the model, and the tools/ directory includes local inference scripts for both SmolLM and SmolVLM.

Editorial conclusion

SmolLM3 is worth evaluating for teams that need a small, fully open model for inference on-device or on modest hardware, particularly if multilingual support across English, French, Spanish, German, Italian, and Portuguese is a requirement. The 3B model runs on consumer GPUs when loaded in standard precision. Engineers who need the highest possible reasoning accuracy at small scale should verify SmolLM3's performance on their specific task before committing, since the README compares it against Llama 3.2 3B and Qwen2.5 3B but does not document task-specific benchmarks in the repository. The Apache-2.0 license and fully open training details make it compatible with commercial use and reproducibility research.

Frequently asked questions

What is SmolLM good for?

SmolLM3 is designed for on-device inference and applications that need a capable language model without the memory requirements of a 7B or larger model. It supports six languages natively, long context up to 128k tokens, and dual-mode reasoning with think and no-think modes.

Who created SmolLM?

SmolLM is developed by Hugging Face, specifically the HuggingFaceTB (Technical Bridge) team. The SmolLM2 technical report is authored under the Hugging Face organization and is available at arXiv:2502.02737v1.

Is SmolLM multilingual?

Yes. SmolLM3 supports six languages: English, French, Spanish, German, Italian, and Portuguese. This multilingual support is built into the 3B model and does not require separate fine-tuning for each language.

Is SmolLM open source?

The huggingface/smollm repository is Apache-2.0 licensed. SmolLM3 is described as fully open: weights, training data mixture, and training configurations are all public. Model weights have their own license on Hugging Face that should be checked separately before commercial deployment.

Official sources

  1. huggingface/smollm on GitHub
  2. Issues
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/huggingface-smollm.svg)](https://hysenlabs.com/projects/huggingface-smollm)