PhoBERT v2 is the checkpoint with more data and a copyleft license, and the fast tokenizer still lives on a fork
PhoBERT: Pre-trained language models for Vietnamese (EMNLP-2020 Findings)
At a glance
- What is it?
- The Vietnamese encoder's repository is four files and two licenses. Everything users actually run is on Hugging Face, and the page points at its own successor for the two constraints it spends most of its length describing.
- Who is it for?
- PhoBERT earns its place in a Vietnamese pipeline in two specific ways. It was the first large-scale monolingual encoder for the language, and it still holds up as a fine-tuning target for part-of-speech tagging, dependency parsing, named-entity recognition and natural language inference.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 67 days ago.
- What is it written in?
- GitHub does not report a main language for this repository.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The repository is four files, and two of them are licenses
The root of the project holds `LICENSE`, `LICENSE_for_PhoBERT_v2`, `README.md` and `README_fairseq.md`. There is no source directory, no scripts and no tests, and no primary language is recorded for the repository at all. Nothing is built from here.
Everything a user runs lives on the Hugging Face hub, where three checkpoints are published under the `vinai` organisation. The default branch is `master` rather than `main`, there are no GitHub releases, and the last push is dated 2026-08-04. The repository's job is documentation, citation guidance and licensing, which is why it can stay small while the model has been in use for years.
That shape has a practical consequence. There is no version history on this repository to read, so anyone asking what changed in a checkpoint has to go to the hub or the paper. The README does make one version check possible: the models table distinguishes v2 from the original pair by name, parameter count and pre-training data, and gives each its own license link.
The checkpoint with more data is the one under a copyleft license
Three checkpoints are published and they do not share a license.
| model | params | max length | pre-training data | license | | --- | --- | --- | --- | --- | | phobert-base-v2 | 135M | 256 | 20GB Wikipedia and News plus 120GB from OSCAR-2301 | GNU Affero GPL v3 | | phobert-base | 135M | 256 | 20GB of Wikipedia and News | MIT | | phobert-large | 370M | 256 | 20GB of Wikipedia and News | MIT |
So v2 is the one trained on six times more text, and v2 is the one under AGPL v3. The original base and large checkpoints remain MIT. That is why there is a second license file in the root rather than a single one covering the project, and the links in the table point at different documents.
For a research group or a commercial product this is the decision, not the parameter count. The MIT checkpoints are the permissive ones, and they are also the ones trained on a fifth of the data. Nothing on the page offers an MIT-licensed v2 or explains whether one is planned.
The fast tokenizer is still installed from a personal fork branch
Installation starts with `pip install transformers`, and then the page goes somewhere unexpected. A slow tokenizer for PhoBERT was merged into the main `transformers` branch, and the fast tokenizer was described as still under discussion at the time of writing, tied to a pull request and an issue comment. To get the fast one, the instruction is to clone a fork:
git clone --single-branch --branch fast_tokenizers_BARTpho_PhoBERT_BERTweet https://github.com/datquocnguyen/transformers.git
cd transformers
pip3 install -e .Three things are worth noticing. The branch lives on an individual's fork, not on the Hugging Face repository. Its name covers three projects at once, BART, PhoBERT and BERTweet, so this is a working branch rather than a released tokenizers version. And nothing on the page says whether the merge has since happened upstream, so a reader arriving today cannot tell from the README whether following these three commands is still necessary.
A separate line adds `pip3 install tokenizers`. The main library also has a TensorFlow 2 route, but in the usage example it is present only as commented-out lines.
All three checkpoints stop at 256 tokens
The models table gives the same maximum length for every checkpoint, 256 tokens, whether base or large and whichever license applies. There is no long-input variant and no sliding-window note.
That number is the practical limit on what this model can be used for. A Vietnamese paragraph or a short document clause fits; a long legal sentence, a review with context, or a multi-turn prompt does not, and truncation has to happen somewhere before the tokenizer sees the text.
The comparison the page makes with its successor is on exactly this axis. BamiBERT is described as supporting an extended context of up to 2,048 tokens, eight times the figure here, and as operating directly on raw input. Combined with the segmentation requirement below, those are the two constraints that shape what code you have to write around PhoBERT, and they are the two the successor addresses.
For anyone reproducing a published baseline, 256 is also a feature: it is the window the reported numbers were produced under, so a longer-context alternative is not a drop-in substitute for a comparison table.
The tokenizer will not segment for you, and the sample text shows it in the underscores
The usage example carries a comment that is the single most important line in the document: input text must already be word-segmented. The sample sentence carries it visually, `Chúng_tôi là những nghiên_cứu_viên .`, with underscores standing in for the spaces a segmenter would have inserted.
The recommended way to get that is the same pipeline used for pre-training. PhoBERT pre-processed its training data with the RDRSegmenter from VnCoreNLP, including Vietnamese tone normalization and both word and sentence segmentation, and the page advises using the same segmenter on raw input for downstream work rather than reaching for something else.
pip install py_vncorenlpThat import downloads the VnCoreNLP components from their original repository and saves them locally, then loads the segmentation component with an explicit annotator list. So the cost is not just a pip install: it is a second model download plus a pass over your input before the tokenizer runs, on every document you process.
The example output is cut off partway through the first token, at `Đại_học Qu`, so the printed form of the remainder is not visible on the page.
Two documented paths, and one of them is in a different file
The table of contents lists two ways to use the models, one with `transformers` and one with `fairseq`. The `transformers` path is written out in full on this page: install, load the model and tokenizer with the `AutoModel` and `AutoTokenizer` classes, encode, wrap the call in `torch.no_grad()`, and read the output, which the comment notes is now a tuple.
The `fairseq` path is not. It occupies two lines and points at `README_fairseq.md`, a separate document in the same repository. The root page therefore documents half of what it advertises, and the fairseq route requires opening a second file that no summary of this one would lead you to.
The choice between them is not argued for anywhere. Both are given as supported, and the page does not say which to prefer for new work or whether one is maintained. Given that the fork-branch tokenizer instruction lives in the `transformers` section, the practical split is that the pip route works out of the box with a slow tokenizer while the fairseq route lives entirely outside the library ecosystem this repository otherwise targets.
The successor removes both constraints this page spends its length explaining
A single italic paragraph points readers at BamiBERT, described as the same group's new BERT-based model for Vietnamese, trained from scratch on a 129 GB corpus of general-domain Vietnamese text for 20 epochs. It is said to address key limitations of PhoBERT, and two of them are named in the numbers that follow.
Context length goes to 2,048 tokens. Input becomes raw, so external word segmentation is eliminated, which removes the RDRSegmenter step, the tone normalization step and the second model download described above. The training corpus is roughly six times the 20 GB behind the original checkpoints.
The reported results are stated in metric counts rather than scores. Across eight Vietnamese benchmarks it achieves the best performance on 11 of 15 metrics and the second-best on three others, and it is described as setting a new state of the art among base-sized Vietnamese encoders with strong cross-domain generalization.
The wording of the recommendation is deliberately mild, that users may also want to use it. But a reader comparing the two specifications directly is looking at a model with a longer window, no preprocessing dependency and more data, which makes this paragraph the most consequential content on the page.
The citation block is an instruction, not a suggestion
The paper reference is included in the repository as an indented citation block: PhoBERT, Pre-trained language models for Vietnamese, by Dat Quoc Nguyen and Anh Tuan Nguyen, in Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1037 to 1042. A link to the anthology version sits above it.
The sentence that follows is not phrased as a request. It says to cite the paper when PhoBERT is used to help produce published results or when it is incorporated into other software, which is a wider trigger than the first clause. If you are benchmarking against published Vietnamese baselines rather than using the model directly, that sentence is the relevant line.
The results the paper is being cited for are four downstream tasks, part-of-speech tagging, dependency parsing, named-entity recognition and natural language inference, on which the page claims new state-of-the-art performance over previous monolingual and multilingual approaches. The architecture itself follows RoBERTa, which in turn optimizes the BERT pre-training procedure.
The four-entry repository means this citation, the license table and the successor note are the three things the project actually asks of you.
Editorial conclusion
PhoBERT earns its place in a Vietnamese pipeline in two specific ways. It was the first large-scale monolingual encoder for the language, and it still holds up as a fine-tuning target for part-of-speech tagging, dependency parsing, named-entity recognition and natural language inference. Where it is dated is on the two axes the ecosystem has since moved on: a 256-token window and a mandatory word-segmentation step in front of the tokenizer. Three things to settle before you build on it. Decide the license per checkpoint, not per repository, because the repository is MIT while phobert-base-v2 is AGPL v3 and ships a second license file for exactly that reason; v2 is the checkpoint with the extra 120GB of OSCAR data and it is also the one with copyleft terms. Read the fast-tokenizer section with suspicion, because it still routes you to clone a personal fork branch rather than install from PyPI, and nothing on the page says whether that has been resolved upstream. And budget the segmentation step, because the tokenizer will not do it for you and the recommended segmenter is a separate download through py_vncorenlp. If any of those three are problems, read the successor note before you commit. BamiBERT, from the same group, trains from scratch on 129 GB for 20 epochs, opens its context to 2,048 tokens, takes raw unsegmented input, and is described here as best on 11 of 15 metrics across eight Vietnamese benchmarks. That is not a minor upgrade; it removes both of this model's sharpest edges. What PhoBERT still has is history and adoption, and for reproducing published Vietnamese baselines that is often the deciding factor.
Frequently asked questions
how to use phobert
Install transformers, load vinai/phobert-base with AutoModel and AutoTokenizer, and remember that the input text must already be word-segmented. For raw text, install py_vncorenlp and run the RDRSegmenter from VnCoreNLP first, which is the same segmenter used to pre-process the training data. A fairseq route exists but is documented in a separate README_fairseq.md file.
What is the license difference between the PhoBERT checkpoints?
The repository is MIT, but the checkpoints differ. vinai/phobert-base and vinai/phobert-large are MIT licensed, while vinai/phobert-base-v2 is GNU Affero GPL v3 and ships with a separate LICENSE_for_PhoBERT_v2 file. The v2 checkpoint is also the one trained on the larger corpus, with 120GB from OSCAR-2301 added to the 20GB of Wikipedia and news text.
Does the phoBERT tokenizer split Vietnamese words automatically?
No. The example code states that input text must already be word-segmented, and the sample sentence shows it with underscores standing in for spaces. The page recommends the RDRSegmenter from VnCoreNLP, which also performs Vietnamese tone normalization, and notes that this is the same segmenter used for pre-training.
How long an input can the phoBERT models handle?
All three published checkpoints list a maximum length of 256 tokens, base or large. BamiBERT, the successor model named on the same page, is described as supporting an extended context of up to 2,048 tokens and operating directly on raw input without external word segmentation.
Why does installing the phoBERT fast tokenizer involve a git clone?
The page states that a slow tokenizer was merged into the main transformers branch while the fast tokenizer was still being discussed, and gives a clone of a branch on an individual fork named fast_tokenizers_BARTpho_PhoBERT_BERTweet as the route to the fast version. Nothing on the page says whether that merge has since happened upstream.
Is phoBERT still a good choice over BamiBERT for Vietnamese?
The page positions BamiBERT as addressing key limitations of PhoBERT, with a 2,048-token context, raw unsegmented input, and training from scratch on a 129 GB corpus for 20 epochs. It reports best performance on 11 of 15 metrics across eight Vietnamese benchmarks. PhoBERT keeps the advantage of matching the 256-token window its published baselines were measured under.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/vinairesearch-phobert)