# MOSS: Fudan's 16B bilingual chat model, its plugin variants, and the hardware you need

> MOSS is an Apache-2.0 code release around a 16B-parameter bilingual dialogue model, with quantized weights, full SFT data and a finetuning script. The catch is memory: FP16 inference needs a 31GB load, and the plugin weights are a separate download.

**OpenMOSS/MOSS** — An open-source, tool-augmented conversational language model from Fudan University

- Repository: https://github.com/OpenMOSS/MOSS
- Website: https://txsun1997.github.io/blogs/moss.html
- Stars: 12,253 · Forks: 1,127
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/openmoss-moss

## What MOSS is, and the problem it addresses

MOSS is a bilingual (Chinese and English) conversational language model released by the OpenMOSS team at Fudan University. The repository holds the MOSS-003 generation of models, the dialogue data used to train them, and the code to deploy and finetune them. The moss-moon series has 16 billion parameters.

The problem it addresses is narrower than "open source ChatGPT". Anyone building a Chinese-language assistant has a small pool of base models to start from, and even fewer with published instruction-tuning data. MOSS ships both: the moss-003-sft-data set, described in the README as roughly 1.1 million conversations built from about 100,000 user inputs collected during the MOSS-002 internal test plus gpt-3.5-turbo, and the plugin-augmented set of roughly 300,000 multi-turn conversations covering four plugins: a search engine, text-to-image, a calculator and an equation solver.

That data release is the part worth attention. Weights alone tell you nothing about what a model was tuned to do; conversation data shows the distribution it was shaped toward, including the usefulness, faithfulness and harmlessness labels the README describes. If your goal is to study or reproduce instruction tuning for Chinese dialogue, this is the material. If your goal is a drop-in assistant, the model card matters more than the data, and that is where the memory numbers bite.

## How the moss-moon-003 variants differ

The release is not one model but a family, and the naming is the map. moss-moon-003-base is the pretrained base, self-supervised on roughly 700B words of Chinese, English and code, with a stated compute budget of about 6.67x10^22 floating point operations. moss-moon-003-sft is that base tuned on the 1.1M conversation set, and the README credits it with instruction following, multi-turn dialogue and refusal of harmful requests. moss-moon-003-sft-plugin is tuned on the conversation data plus the plugin data, adding the four plugin capabilities on top of the sft model.

Each of those two has a 4-bit and an 8-bit build, so the practical choice is six downloadable checkpoints. The README also lists a preference model, moss-moon-003-pm, and the final preference-trained models moss-moon-003 and moss-moon-003-plugin, all marked as coming soon. That distinction matters: the strongest models in the family are announced but not yet available, and the repository does not give a date.

Pick by capability, not by size. If you need tool calls, the plugin variants are the only ones trained for them; the README describes the plugin data as the source of that ability, so the plain sft checkpoint has no reason to have it. If you only need chat, the sft variants are the smaller commitment.

## Hardware requirements and the memory table

The README publishes a table for batch size 1, and it is the first thing to read. Loading the model in FP16 takes 31GB; finishing one round of conversation is estimated at 42GB, and reaching the maximum conversation length of 2048 tokens at 81GB. Int8 drops those to 16GB, 24GB and 46GB. Int4 drops them to 7.8GB, 12GB and 26GB.

Read the middle column, not the first. A 24GB card that loads an int8 model comfortably will not necessarily hold a long conversation, because the KV cache grows with context. The 46GB figure for a full 2048-token int8 conversation is the number that decides whether a single 3090 is enough for your workload. The README states that FP16 runs on a single A100/A800 or two 3090 cards, and that int4 and int8 run on a single 3090.

One constraint is easy to miss: the README states that quantized models do not currently support model parallelism. If you planned to split an int4 model across two GPUs to gain headroom, that path is closed, and the 7.8GB load figure is for a single device. Plan capacity against the conversation column and the no-parallelism rule together.

## Installing MOSS and running a first conversation

The README gives a three-step setup. Clone the repository, create a Python 3.8 conda environment, and install the pinned dependencies. The requirements file pins torch to 1.13.1 and transformers to 4.25.1, and the README says versions of torch and transformers below the recommended ones are not advised, so treat the pins as the supported combination rather than a suggestion.

```bash
git clone https://github.com/OpenMOSS/MOSS.git
cd MOSS
conda create --name moss python=3.8
conda activate moss
pip install -r requirements.txt
```

The dependencies include triton, and the README states that triton currently supports Linux and WSL only, not Windows or macOS. On a Mac or a Windows machine without WSL, the install will not complete as written. The README says to wait for a later update; it does not offer an alternative path.

The README's usage example loads the tokenizer and model through transformers with trust_remote_code enabled, using the sft checkpoint. The snippet in the README is truncated at the point where the model is moved to a device, so the shape of the call is clear even though the full script is not reproduced there.

```python
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("OpenMOSS-Team/moss-moon-003-sft", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("OpenMOSS-Team/moss-moon-003-sft", trust_remote_code=True)
```

For an interactive session, the repository also ships moss_cli_demo.py, moss_web_demo_gradio.py and moss_web_demo_streamlit.py, and a Jittor path with moss_cli_demo_jittor.py and a models_jittor directory. The README does not document the command-line flags or environment variables for those demo scripts, so read the files themselves before assuming an interface. If you want the plugin behaviour, you load a plugin checkpoint instead, and the README points to MOSS WebSearchTool as the separate deployment scheme for the search plugin.

## Finetuning and the data that ships with it

The repository includes finetune_moss.py, a configs directory and the SFT_data tree. The README lists the finetuning section with software dependencies and usage, and the data is not a sample: the moss-003-sft-data set is described as fully open, and the plugin conversation data is published on Hugging Face under moss-003-sft-data.

This is the strongest argument for MOSS over a weights-only release. You can inspect the conversation distribution your finetune will be measured against, and the harmlessness and usefulness labels are part of the data rather than an undocumented filter. For a team building a domain-specific Chinese assistant, starting from a model whose tuning data you can read is a real advantage over starting from a model card alone.

The cost is that finetuning a 16B model is not a small job, and the README does not publish a finetuning memory table the way it does for inference. Nothing in the repository states the GPU memory or time required to run finetune_moss.py, so treat that as unmeasured until you check the script and configs yourself.

## Limitations, and when MOSS is the wrong tool

The README states its own limitation plainly: because of the small parameter count and the autoregressive generation paradigm, MOSS can still produce misleading replies containing factual errors, or harmful content containing bias or discrimination, and users should judge its output carefully and not spread harmful content it generates.

That is not boilerplate. A 16B model tuned on 1.1M conversations sits well below the frontier, and the README's own examples of code generation and Chinese conversation are illustrative screenshots, not evaluations. There is no benchmark table in the repository, and no accuracy numbers for the plugins. If your decision depends on measured quality, this repository does not give you one.

Three cases where MOSS is the wrong tool. First, anything needing a small local footprint: the smallest supported configuration is still a 3090 with an int4 model, and the README's own table puts a full-length conversation at 26GB even then. Second, Windows or macOS without WSL, since triton is Linux and WSL only. Third, production systems that need predictable tool execution: the plugin checkpoints are separate downloads, the search plugin has its own deployment repository, and the README does not document what the model does when a plugin call fails or returns nothing.

A fourth point is licensing rather than capability. The code is Apache-2.0, but the README's badges show CC BY-NC 4.0 for the data and GNU AGPL 3.0 for the model, with separate LICENSE, DATA_LICENSE and MODEL_LICENSE files at the repository root. Those are three different terms on three artifacts in one repository. The non-commercial data licence in particular deserves a read before you plan a commercial finetune on moss-003-sft-data. This is not legal advice; read the files and, if the stakes justify it, get proper review.

## How MOSS compares with LLaMA-style releases

The obvious alternative is a LLaMA-family base model with a community finetune, and the difference is not quality, it is what gets published. LLaMA-derived releases typically distribute weights and a recipe, and the instruction data is often synthetic or withheld. MOSS distributes weights, the SFT conversations, the plugin conversations, the preference-training plan and the finetuning script in one repository, under a Chinese-first bilingual design.

That makes MOSS the better starting point when your target language is Chinese and you want to see the tuning distribution, or when you are studying how plugin data changes a model's behaviour, since the plugin and non-plugin checkpoints are otherwise identical in lineage. It makes MOSS the worse starting point when you need the largest ecosystem of adapters, quantized community builds and serving integrations, because a 16B bilingual model with pinned torch 1.13.1 and transformers 4.25.1 has a narrower set of tools around it. The pinning is itself a signal: this is a research release with a tested combination, not a library chasing the latest transformers.

Within the MOSS project there is also a Jittor path alongside the PyTorch one, with a separate models_jittor directory and CLI demo. If your stack is Jittor, that is an option the LLaMA ecosystem does not offer; if it is not, the PyTorch path is the one the README documents in detail.

## Conclusion

Adopt MOSS if you want a bilingual 16B dialogue model whose training conversations are published alongside the weights, and you have at least one 3090 or an A100/A800 to run it. Do not adopt it if you need a small local assistant, a production SLA, or plugin behaviour you have not measured yourself, because the plugin weights ship as a separate model and the README does not document failure handling. Verify three things first: that your GPU memory matches the table for your chosen precision, that the int4/int8 build suits your deployment since quantized models do not support model parallelism, and that MODEL_LICENSE and DATA_LICENSE fit your use, because the Apache-2.0 badge covers the code only.

## FAQ

### What is MOSS?

MOSS is an open source bilingual Chinese and English dialogue language model from the OpenMOSS team at Fudan University. The repository provides the MOSS-003 series models, the conversation data used to train them, and deployment and finetuning code.

### Is there an open-source voice generator available?

MOSS itself is a text dialogue model, not a voice generator; the README describes text conversation, code generation and four plugins (search, text-to-image, calculator, equation solver). The repository's related projects list includes MOSS-TTS, which is a separate project from this one.

### What is the best TTS open source?

The MOSS repository covers a text dialogue model and its training data, and its related projects list points to MOSS-TTS as a separate repository, but no text-to-speech comparison or evaluation is given in this repository.

### What is text to spoken dialogue generation and how does it work?

The MOSS repository does not describe spoken dialogue generation. The moss-moon models generate text: the README credits moss-moon-003-sft with instruction following and multi-turn dialogue, and moss-moon-003-sft-plugin with four plugins covering search, text-to-image, a calculator and an equation solver.

## Sources

- [Issues](https://github.com/OpenMOSS/MOSS/issues)
- [License: Apache-2.0](https://github.com/OpenMOSS/MOSS/blob/main/LICENSE)
- [OpenMOSS/MOSS on GitHub](https://github.com/OpenMOSS/MOSS)
- [Project website](https://txsun1997.github.io/blogs/moss.html)
- [README](https://github.com/OpenMOSS/MOSS/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/openmoss-moss
