FastChat: what the two documented install commands leave out
An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
At a glance
- What is it?
- lm-sys/FastChat is the Apache-2.0 platform that serves open chat models and powers Chatbot Arena, with Vicuna and MT-Bench built on it. Its two documented install paths pull only the model_worker and webui extras, so the evaluation, training, and dev dependencies are never installed by either command.
- Who is it for?
- FastChat fits a researcher who wants a local, OpenAI-compatible server for Vicuna or LongChat, or who wants to reproduce the MT-Bench side of the pipeline. It does not fit someone who needs the judge or fine-tuning stack out of the box, because the documented install never adds those extras, and it does not fit anyone who needs a released, pinned package matching the current source.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 153 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Two install paths, and one of them edits your checkout
The package name on PyPI is fschat, and the short path installs it with the two extras most people need:
pip3 install "fschat[model_worker,webui]"The second path clones the repository and installs the same extras in editable mode:
git clone https://github.com/lm-sys/FastChat.git
cd FastChatpip3 install --upgrade pip # enable PEP 660 support
pip3 install -e ".[model_worker,webui]"The upgrade line is not decoration. Editable installs go through PEP 660, so an old pip on the machine is the first thing that fails. Mac users are told to install two build tools before that step:
brew install rust cmakeThe consequence for you is a real fork in behaviour. Editable mode means the fastchat package you import is whatever sits on the branch you cloned, including any local edit you make, while the pip form is frozen at whatever version was published.
Three extras are named in pyproject and installed by neither command
pyproject.toml defines five optional extras, and the two documented install commands name two of them. model_worker carries accelerate>=0.21, peft, sentencepiece, torch, transformers>=4.31.0, protobuf, openai, and anthropic. webui carries gradio>=4.10, plotly, and scipy. The base dependency list is the serving stack: aiohttp, fastapi, httpx, markdown2[all], nh3, numpy, prompt_toolkit>=3.0.0, pydantic<3,>=2.0.0, pydantic-settings, psutil, requests, rich>=10.0.0, shortuuid, tiktoken, and uvicorn.
Three extras are never installed by either command: train with einops, flash-attn>=2.0, and wandb, llm_judge with openai<1, anthropic>=0.3, and ray, and dev with black==23.3.0 and pylint==2.8.2.
The consequence is that a bare install serves an API but cannot load a model, and that following the quick start to the end still leaves the judge uninstalled even though the evaluation code is one of the two advertised core features. The llm_judge extra also holds openai below 1.0, a boundary you inherit silently.
14GB for 7B and 28GB for 13B decide which command you can copy
The single GPU command is the one most people paste first:
python3 -m fastchat.serve.cli --model-path lmsys/vicuna-7b-v1.5It needs around 14GB of GPU memory for Vicuna-7B and 28GB for Vicuna-13B. The CPU only path needs around 30GB of CPU memory for Vicuna-7B and around 60GB for Vicuna-13B, so a workstation with 32GB of RAM is inside the 7B figure and outside the 13B one. The same flag serves both cases, because --model-path accepts a local folder or a Hugging Face repo name.
What the CLI does not offer is any way around those numbers. No quantization switch, no CPU offload flag, and no device list appear in the documented options, and a reader short on memory is pointed to a separate Not Enough Memory section rather than given a flag. The consequence is that your hardware, not your code, picks the model, and a 24GB card that could hold a quantized 13B in another stack has nothing to paste here.
`--max-gpu-memory` exists because auto device mapping does not balance
Model parallelism aggregates memory across GPUs on one machine:
python3 -m fastchat.serve.cli --model-path lmsys/vicuna-7b-v1.5 --num-gpus 2The documented reason for a second flag is that the auto device mapping strategy in huggingface/transformers does not perfectly balance the memory allocation across multiple GPUs. So you can cap how much each GPU holds for model weights:
python3 -m fastchat.serve.cli --model-path lmsys/vicuna-7b-v1.5 --num-gpus 2 --max-gpu-memory 8GiBLeaving that headroom is the point, because the freed space goes to activations, which is what buys longer context lengths and larger batch sizes.
The consequence is that multi GPU serving is manual work, not auto configuration. One example value, 8GiB, is given and no per model guidance follows, and the flag caps weight memory only, with nothing said about how the remaining allocations behave at long context.
First launch is a download into the cache folder in your home directory
The chat commands are also the download commands. Passing a Hugging Face repo name to --model-path pulls the weights on first use, and the files land in a .cache folder in the user's home directory, in the shape ~/.cache/huggingface/hub/<model_name>. Nothing prompts for a target directory.
One version constraint is stated and easy to miss: transformers>=4.31 is required for the 16K versions. That pin lives in the model_worker extra rather than the base requirements, so it arrives only when you install that extra, and a global environment with an older transformers will fail on the 16K rows while the 7B and 13B rows of the same table still work.
The consequence is that a first run needs network access, home directory space, and a matching transformers, and the CLI documents no offline mode, no mirror setting, and no way to relocate that cache. On a machine with a small home volume or no egress, the failure lands at the moment you expected a prompt.
The 33B row points at v1.3 while every other Vicuna row is v1.5
The weights table lists five Vicuna chat commands. The 7B row uses lmsys/vicuna-7b-v1.5, the 7B-16k row uses lmsys/vicuna-7b-v1.5-16k, the 13B row uses lmsys/vicuna-13b-v1.5, the 13B-16k row uses lmsys/vicuna-13b-v1.5-16k, and the 33B row uses lmsys/vicuna-33b-v1.3.
Four rows sit on v1.5 and one sits on v1.3, with no note in the table explaining the split. The pointer for the history is a separate file, docs/vicuna_weights_version.md, which the table itself links as the place for all versions and their differences.
The consequence is concrete. A reader who assumes the table is one release generation pulls a 33B model from a different point in the project's history than the 13B model they tested an hour earlier, and the two do not share a version label, a context length story, or a documented behaviour. For evaluation work that difference is the kind of variable you only notice after the numbers come out wrong.
Apache-2.0 covers the code, and the weights are under a different license
The repository itself carries Apache-2.0, and pyproject repeats it in the classifier list. The weights do not travel with it. Vicuna is based on Llama 2 and should be used under Llama's model license, which is a separate document living in the Llama repository, and the same split applies to the other two models the project released, LongChat and FastChat-T5, whose chat commands point at lmsys/longchat-7b-32k-v1.5 and lmsys/fastchat-t5-3b-v1.0.
The consequence is that a clean Apache-2.0 grant on the serving code says nothing about the model you are serving, the weights arrive from Hugging Face repos under their own terms, and anyone who ships a service built on Vicuna has an obligation that the repository license header does not describe. The README states the obligation in one sentence and does not restate the conditions, so the actual terms are something you have to go and read for yourself before you deploy.
pyproject says 0.2.36, and the newest release is from February 2024
The version field in pyproject.toml reads 0.2.36, and that matches the most recent GitHub release, v0.2.36, dated 2024-02-11. The two releases before it are v0.2.35 from 2024-01-17 and v0.2.34 from 2023-12-09. The repository is not archived and the last push is dated 2026-05-01, so commits have continued for more than two years without a new tag.
What that gap does to a reader is specific. pip3 install fschat gives you the February 2024 build, and any fix that landed on main after that is in your clone but not in your environment, which is exactly the split the editable install path creates for anyone who mixes the two. The version in pyproject does not move, so a project directory cannot tell you which state it is in. The dev extra compounds this by pinning black==23.3.0 and pylint==2.8.2, so the format.sh script and the .pylintrc at the repository root only agree with your global tools if you install that extra.
Editorial conclusion
FastChat fits a researcher who wants a local, OpenAI-compatible server for Vicuna or LongChat, or who wants to reproduce the MT-Bench side of the pipeline. It does not fit someone who needs the judge or fine-tuning stack out of the box, because the documented install never adds those extras, and it does not fit anyone who needs a released, pinned package matching the current source. Before you commit, check the version in pyproject.toml against the newest GitHub release, read the Llama 2 license rather than the repository's Apache-2.0, and confirm your GPU memory against the published figures for the size you plan to run.
Frequently asked questions
how to install fastchat
Either pip3 install "fschat[model_worker,webui]" for the published package, or clone the repository and run pip3 install -e ".[model_worker,webui]" after upgrading pip, with brew install rust cmake first on Mac. Both commands add only the model_worker and webui extras.
what is fast chat
FastChat is an open platform for training, serving, and evaluating large language model based chatbots. It powers Chatbot Arena, which the README says has served over 10 million chat requests for 70+ LLMs and collected over 1.5M human votes.
Does the FastChat pip install bring in the evaluation dependencies?
No. The two documented install commands name only the model_worker and webui extras. The llm_judge extra, which carries openai<1, anthropic>=0.3, and ray, and the train extra, which carries einops, flash-attn>=2.0, and wandb, are defined in pyproject.toml and installed by neither command.
Where does FastChat put the model weights it downloads?
Downloaded weights are stored in a .cache folder in the user's home folder, in the shape ~/.cache/huggingface/hub/<model_name>. The first chat command triggers the download from the Hugging Face repo named in --model-path, which accepts either a repo name or a local folder.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lm-sys-fastchat)