Model or dataset
bilibili/Index-1.9B avatar
bilibili/Index-1.9B

Index-1.9B: bilibili's 1.9B-Parameter Multilingual Model and Its 32K Variant

A lightweight multilingual LLM

1,033 stars53 forksPythonApache-2.0

At a glance

What is it?
Index-1.9B is a small Chinese-first language model family from bilibili, with chat, role-playing, a control variant and a 32K long-context build. It is cheap to run and easy to load, but the 32K version only starts through one specific demo script.
Who is it for?
Adopt Index-1.9B if you want a small, Chinese-first model you can run on CPU or a single consumer GPU and you are willing to read the demo scripts instead of the README. Skip it if you need an English-first general assistant or a documented deployment path, because the repository gives you demos, not a serving story.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Index-1.9B is, and who the five checkpoints are for

The Index-1.9B series is described in the README as a lightweight version of bilibili's Index model line. The base model has 1.9 billion non-embedding parameters and was pre-trained on 2.8T tokens of mainly Chinese and English corpus. That size is the whole point: it fits on modest hardware, and the training data skews Chinese, which is where the model is meant to be used.

Five checkpoints are published. Index-1.9B base is the pre-trained model. Index-1.9B pure is a control version with the same parameters and training strategy but with instruction-related data filtered out of the corpus, so the authors can measure what instruction data contributes to benchmark scores. Index-1.9B chat is aligned with SFT and DPO on top of the base; the README claims the pre-training mix included a lot of internet community corpus, which it says gives the chat model more interesting conversational behavior and stronger multilingual translation, especially for East Asian languages. Index-1.9B character adds RAG on top of SFT and DPO for few-shot role-playing. Index-1.9B-32K extends context length to 32K, which the README describes as handling documents of over 35,000 words in one pass.

If you are building a Chinese-language assistant on a single GPU, the chat checkpoint is the one you want. The pure checkpoint exists for researchers studying instruction tuning, not for applications. The character checkpoint is a narrow tool: it is for role-play customization, and the repository ships a roleplay/ directory to match.

How the model family is put together

The architecture is not documented in the README beyond parameter counts, so the useful information is in the repository layout. There are four top-level code directories: demo/ for inference entry points, evaluate/ for benchmarking, finetune/ for training, and roleplay/ for the character model. The evaluation code is based on OpenCompass with compatibility modifications, and the README points at the evaluate/ folder for details.

The data flow for a normal chat session is the standard transformers pipeline: a tokenizer and a text-generation pipeline are loaded from a local model directory, the conversation is assembled as a list of role and content dictionaries with a system message first, and the pipeline returns generated text. The README's example passes max_new_tokens, top_k, top_p, temperature, repetition_penalty and do_sample directly to the generator call.

The 32K variant is the exception. The README states in bold, twice, that Index-1.9B-32K can only be launched using demo/cli_long_text_demo.py. That means the long-context path is not the same code path as the standard chat demo, and the repository contains a separate technical report, Index-1.9B-32K_Long_Context_Technical_Report.md, alongside its Chinese counterpart. If you plan to use the 32K model as a drop-in replacement for the chat model in your own script, the README gives you no reason to think that works.

Installing Index-1.9B and running your first chat

The repository is cloned and dependencies installed with pip. requirements.txt pins two packages: gradio==4.29.0 and transformers==4.39.2. Note that gradio is only needed for the web demo, and the README separately tells you to install flask==2.2.5 if you want the OpenAI-compatible API demo.

bash
git clone https://github.com/bilibili/Index-1.9B
cd Index-1.9B
pip install -r requirements.txt

Download a checkpoint from HuggingFace or ModelScope (both links are in the README's model table), then load it with transformers. The README warns that the model directory must not contain a dot and can be replaced with an underscore, which is why the example path uses Index-1.9B-Chat rather than a literal repository-style name.

python
import argparse
from transformers import AutoTokenizer, pipeline

parser = argparse.ArgumentParser()
parser.add_argument('--model_path', default="./IndexTeam/Index-1.9B-Chat/", type=str, help="")
parser.add_argument('--device', default="cpu", type=str, help="")
args = parser.parse_args()

tokenizer = AutoTokenizer.from_pretrained(args.model_path, trust_remote_code=True)
generator = pipeline("text-generation", model=args.model_path, tokenizer=tokenizer, trust_remote_code=True, device=args.device)

The device argument accepts "cpu", "cuda" or "mps" for Apple silicon, so you can start on CPU before moving to a GPU. For an interactive session, the terminal demo takes a model path directly:

bash
python demo/cli_demo.py  --model_path='/path/to/model/'

For a browser UI, install gradio and pass a port and model path to demo/web_demo.py; the README shows --port='port' --model_path='/path/to/model/'. For an OpenAI-style endpoint there is demo/openai_demo.py, which needs flask. The README's own example for that script is truncated mid-path, so read the file rather than copying the snippet.

The 32K model has one documented entry point, and that is a real constraint

The clearest limitation in the repository is stated by the project itself: Index-1.9B-32K can only be launched using demo/cli_long_text_demo.py. Every other demo in the repository is written for the standard chat model. If your application needs a library call rather than a terminal session, the README does not tell you how to reach the 32K model, and it does not describe how the long-context attention is implemented outside the separate technical report file.

There is a second, quieter constraint. The README does not document a serving path. There is no vLLM or TGI integration mentioned, no quantization instructions in the README itself, and no rollback or versioning guidance. The GGUF adaptation for llamacpp and Ollama is mentioned in the updates section and points to a separate HuggingFace repository, Index-1.9B-Chat-GGUF, which is where that path lives.

Finally, the benchmark table is a snapshot, not a promise. Index-1.9B's average score of 64.92 sits below Qwen2-1.5B's reported 65.17 and above Qwen1.5-1.8B's 58.96 in the same table, but those comparison numbers come from other projects' reports as cited by this README. Treat the table as the authors' own evaluation under modified OpenCompass, not as an independent result.

How it compares to Qwen and MiniCPM at the same scale

The obvious alternatives are the other sub-2B models in the README's own table. Qwen1.5-1.8B and Qwen2-1.5B are the closest match on size and Chinese capability; MiniCPM-2.4B-SFT is slightly larger and scores higher on the English average in that table. The difference in approach matters more than the scores.

Index-1.9B is built around a Chinese-first corpus with a deliberate injection of internet community text, and the README ties the chat model's character and East Asian translation ability to that choice. The Qwen line is a general multilingual family with a much wider set of released sizes and a longer history of tooling integrations. If you need an English-first assistant, or you want a model with an established vLLM or llama.cpp story, Qwen is the safer pick on ecosystem grounds alone, regardless of what the table says.

Where Index-1.9B differentiates is the character checkpoint with RAG-based few-shot role-playing, and the pure checkpoint for instruction-tuning research. Neither has an obvious equivalent in the comparison set. If your project is Chinese role-play or an ablation study on instruction data, the alternatives in the table do not cover the same ground.

Licence, maintenance and the cost of upgrading

The repository carries two licence files: LICENSE and INDEX_MODEL_LICENSE. The GitHub repository metadata lists Apache-2.0, but the presence of a separate model licence means the weights may be governed by terms that differ from the code. The README does not summarize either licence, so read INDEX_MODEL_LICENSE before you ship anything built on these checkpoints. That is a factual observation about the file layout, not legal advice.

The last push to the repository was on 2026-08-21, and the repository is not archived. There are no retrieved releases, so upgrades arrive as commits on main rather than tagged versions. That has a practical consequence: if you pin transformers==4.39.2 as requirements.txt does, you are pinned to a version from the project's own testing, and you should expect to re-check the demo scripts after any pull, because the 32K entry point is a script rather than a stable API. The finetune/ directory exists for training your own variants, which is the more controlled upgrade path if you need one.

Editorial conclusion

Adopt Index-1.9B if you want a small, Chinese-first model you can run on CPU or a single consumer GPU and you are willing to read the demo scripts instead of the README. Skip it if you need an English-first general assistant or a documented deployment path, because the repository gives you demos, not a serving story. Before you commit, check the INDEX_MODEL_LICENSE file against your intended use and confirm that Index-1.9B-32K actually launches through demo/cli_long_text_demo.py on your hardware, since the README says that is the only supported entry point.

Frequently asked questions

Which Index-1.9B checkpoint should I use for a chatbot?

Index-1.9B chat, which the README describes as aligned with SFT and DPO on top of the base model. The base and pure checkpoints are pre-trained and control models rather than dialogue models.

Can I run Index-1.9B-32K with the normal chat demo?

No. The README states twice that Index-1.9B-32K can only be launched using demo/cli_long_text_demo.py, so the other demos in the repository are not supported entry points for that checkpoint.

What are the dependencies for Index-1.9B?

requirements.txt pins gradio==4.29.0 and transformers==4.39.2. The README separately notes that the OpenAI-style API demo depends on flask==2.2.5, which is not in that file.

Does Index-1.9B work with Ollama or llamacpp?

The README's updates section says the model was adapted to llamacpp and Ollama and points to a separate Index-1.9B-Chat-GGUF repository on HuggingFace. The GGUF files are not part of this repository.

Official sources

  1. bilibili/Index-1.9B on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/bilibili-index-1-9b.svg)](https://hysenlabs.com/projects/bilibili-index-1-9b)