Model or dataset
facebookresearch/MobileLLM avatar
facebookresearch/MobileLLM

MobileLLM: training sub-billion parameter language models for on-device use

MobileLLM Optimizing Sub-billion Parameter Language Models for On-Device Use Cases. In ICML 2024.

1,466 stars94 forksPythonNOASSERTION

At a glance

What is it?
MobileLLM is Meta's ICML 2024 training code for sub-billion parameter language models. It ships a torchrun pretraining script, configs from 125M to 1.5B, and an evaluation script, but no packaged install and no documented rollback path.
Who is it for?
Adopt MobileLLM if you are a research or platform team with multi-GPU nodes who wants to pretrain or reproduce sub-billion parameter models and can supply your own tokenized data in the expected per-node jsonl layout. Do not adopt it if you want a pip-installable inference library or an on-device runtime: the repository is training and evaluation code, and it does not ship a mobile deployment path.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 153 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem MobileLLM targets: good small models, not smaller big ones

Most open language model repositories assume you have a cluster and a trillion tokens. MobileLLM starts from the opposite constraint. The README states the work considers multiple design factors to obtain high-quality LLMs with fewer than a billion parameters, and the repository is the training code behind that paper, published at ICML 2024. The intended user is someone who needs a model small enough to plausibly run on a device but strong enough to be useful on zero-shot commonsense reasoning, and who is willing to train it rather than download a finished product. The repository itself is Python, with configs/ holding per-size model definitions, pretrain.py and pretrain.sh as the training entry point, and eval.py plus eval.sh for evaluation on Wiki. The README also points to HuggingFace collections for the released MobileLLM models and for the later MobileLLM-R1 and MobileLLM-R1.5 releases, so the split is clear: this repo is where you train, HuggingFace is where you download.

What the architecture actually changes: SwiGLU, depth, tied embeddings, GQA

The README names four ingredients: SwiGLU activation, deep and thin architectures, embedding sharing, and grouped-query attention. Those are not independent decorations. Depth over width is the claim that a thin stack buys more quality per parameter at this scale; embedding sharing removes a separate output projection, which matters proportionally more when the vocabulary matrix is a large share of a 125M model; grouped-query attention cuts key-value head count, which is the part that pays off at inference time on memory-bandwidth-limited hardware. The README reports MobileLLM-125M and MobileLLM-350M at a 2.7% and 4.3% accuracy boost over preceding 125M and 350M SoTA models on zero-shot commonsense reasoning, and the tables show the per-task numbers behind those averages, for example MobileLLM-125M at 46.3 average against OPT-125M at 42.6 and GPT-neo-125M at 42.9. The README also states that the updated version scales the design philosophy to MobileLLM-600M, 1B and 1.5B. Read the tables as the authors' reported results, not as an independent reproduction.

Installing and running a first pretraining job

There is no package to install. The README gives requirements as python 3.9 and pytorch >= 2.0, with dependencies from requirement.txt, and the repository layout confirms requirement.txt sits at the top level next to pretrain.sh. The README's Step 1 lists the install as pip install -r requirement.txt.

Data comes next, and this is the step that trips people up. The README says to divide a tokenized dataset, or tokenize your own, and distribute it across the total number of training nodes, where each node comprises 1x8 GPUs. Each line of a jsonl file is a key-value pair of tokenized data, and the README gives the shape as {"token_ids": [1,2,3,4,...]}. The directory structure is one numbered folder per node under a base path, each holding one or more jsonl files. The README notes the training code is compatible with the data pre-processing method in LLM360/amber-data-prep, which is the practical route if you would rather not write a tokenizer pipeline yourself.

Then edit pretrain.sh: set --train_data_local_path to the pre-processed data from the previous step and --input_model_filename to ./configs/{model_size}/. The README states the script initiates training on a 1x8 node setup using torchrun, and that it can be modified to adjust --nnodes and other settings for multi-node configurations such as slurm or torchx. The README's own run instruction is bash pretrain.sh.

The learning rate in the script is for 1x8 nodes with a batch size of 32, and the README is explicit that increasing nodes or batch size requires increasing the learning rate linearly. For evaluation, download the models, update the checkpoint path in eval.sh, and run bash eval.sh. The README does not document what output files training writes or where checkpoints land.

The learning-rate rule and the data layout are the two sharp edges

The linear learning-rate scaling instruction is the single most consequential line in the README, and it is one sentence long. If you move from 1x8 nodes to 4x8 and keep the script's learning rate, you are running a configuration the authors did not describe. The same applies in reverse: shrinking the batch without shrinking the learning rate puts you off the stated recipe. The data layout is the other edge. The numbered directories must match the total node count, so a dataset split for 4 nodes will not silently work on 8; you re-split or you accept the mismatch. There is no schema validator mentioned, so a malformed jsonl line surfaces as a training failure rather than a clear data error. And the README does not document checkpoint resumption, so an interrupted run is a question you have to answer from pretrain.py rather than from the documentation.

Training cost: the real gate on whether this is for you

The README publishes a cost table for training on 1T tokens with 32 NVIDIA A100 80G GPUs: roughly 3 days for 125M, 6 for 350M, 8 for 600M, 12 for 1B, and 18 for 1.5B. Those are the authors' figures for one specific hardware and token budget. They are useful precisely because they are unglamorous: even the smallest model is a multi-day, multi-GPU commitment, and the 1.5B run is close to three weeks of a 32-GPU allocation. If your interest is in fine-tuning an already released checkpoint rather than pretraining from scratch, this repository is the wrong tool, because the README frames it as training code and points at HuggingFace for the models. If your interest is in the newer reasoning-oriented line, the README's news items point to MobileLLM-R1 and MobileLLM-R1.5 as separate releases with their own HuggingFace collections, so check which generation you actually want before you spend time on this repository's scripts.

How it compares with a config-driven training framework

A framework such as Megatron-LM or LitGPT takes the opposite approach: it exposes a broad configuration surface and expects you to assemble a model from components. MobileLLM is narrower. Configs live under configs/{model_size}/, the training entry point is a shell script you edit, and the architecture choices (SwiGLU, deep and thin, embedding sharing, grouped-query attention) are baked into what the configs describe rather than offered as a menu. The trade-off is real in both directions. You get a recipe that the authors actually ran, with published zero-shot numbers for each size and a stated hardware cost, which is more than many research repositories provide. You give up the flexibility of swapping attention variants or activation functions without reading pretrain.py. If you need a training framework that supports arbitrary architectures, MobileLLM is the wrong tool. If you need the specific sub-billion recipe and its reported baselines, the narrower surface is the point.

Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-04-30. The README's news list runs through January 2026, when MobileLLM-R1 was accepted to ICLR 2026, and November 2025, when MobileLLM-R1.5 was released. So the project line is still moving, but the movement is mostly in new model releases rather than in this training repository's scripts. Plan for that: the code you adopt today may be the older generation of a family whose newer members live elsewhere. There are no retrieved releases, so there is no versioned artifact to pin against and no changelog to read before upgrading. On licensing, the repository carries a LICENSE file at the top level, but the metadata reports the licence as NOASSERTION, meaning the licence could not be identified automatically. Read LICENSE yourself and get your own advice before using the code or the models commercially; nothing here should be read as legal guidance. The upgrade cost is mostly re-reading pretrain.sh and the configs, because there is no dependency manifest beyond requirement.txt and no documented migration path between generations.

Editorial conclusion

Adopt MobileLLM if you are a research or platform team with multi-GPU nodes who wants to pretrain or reproduce sub-billion parameter models and can supply your own tokenized data in the expected per-node jsonl layout. Do not adopt it if you want a pip-installable inference library or an on-device runtime: the repository is training and evaluation code, and it does not ship a mobile deployment path. Before committing GPU time, verify the LICENSE file directly, confirm that configs/{model_size}/ contains the size you intend to train, and check that pretrain.sh points at a --train_data_local_path whose directory names match your node count.

Frequently asked questions

What does LLM mean in texts?

In this repository the term refers to a large language model, and MobileLLM is training code for sub-billion parameter ones. The README frames the work around obtaining high-quality LLMs with fewer than a billion parameters for on-device use cases.

Is LLM the same as AI?

The README does not discuss that distinction. It treats an LLM as a language model trained on tokenized text, and the repository's data format is a jsonl line holding a key-value pair of tokenized data such as {"token_ids": [1,2,3,4,...]}.

What does LLM mean in chatting?

The README does not describe chat behaviour. The models here are evaluated on zero-shot commonsense reasoning tasks such as arc_easy, arc_challenge, boolq, piqa, siqa, hellaswag, obqa and winogrande, not on conversation.

Official sources

  1. facebookresearch/MobileLLM on GitHub
  2. Issues
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/facebookresearch-mobilellm.svg)](https://hysenlabs.com/projects/facebookresearch-mobilellm)