# TextGen: the portable build, the full install, and what each one can load

> Two distributions, five backends, and one Jinja2 templating pass: how TextGen decides what you are allowed to run, where the documented setup stops, and what the privacy claim actually covers. For developers choosing between the portable folder, the venv route, and the one-click scripts.

**oobabooga/textgen** — Open-source desktop app for local LLMs. Text, vision, tool-calling, OpenAI/Anthropic-compatible API. 100% private.

- Repository: https://github.com/oobabooga/textgen
- Stars: 47,722 · Forks: 5,991
- Language: Python
- License: AGPL-3.0
- Published: 2026-08-17 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/oobabooga-textgen

## The portable build and the full install are two different products

The portable build is a folder with every dependency included, offered for Linux, Windows, and macOS with CUDA, Vulkan, ROCm, and CPU-only options, and it is what you get by downloading, unzipping, and double-clicking `textgen`. The full installation exists for what the portable build does not carry: ExLlamaV3, Transformers, training, image generation, and extensions such as TTS, voice input, and translation. It asks for roughly 10GB of disk and downloads PyTorch. The split decides which models you can load at all. A GGUF file goes into `user_data/models` and the interface detects it automatically. A multi-file model, a 16-bit Transformers model or an EXL3 model, has to sit in a subfolder inside that same directory:

```
textgen
└── user_data
    └── models
        └── Qwen_Qwen3-8B
            ├── config.json
            ├── generation_config.json
            ├── model-00001-of-00004.safetensors
            ├── ...
            ├── tokenizer_config.json
            └── tokenizer.json
```

and those formats require the full installation, not the portable build. There is no equivalent route through a portable copy. Two dates are worth holding on to. The newest release is v4.9, dated 2026-05-20, ahead of v4.8 and v4.7.3, while the last push to the default branch is dated 2026-08-17, so the folder you download from the releases page is not the head of the branch.

## Three entry points are documented, and two files in the tree are not

The quickest documented route is a clone and a virtual environment, and it is the one the project writes out command by command:

```bash
# Clone repository
git clone https://github.com/oobabooga/textgen
cd textgen

# Create virtual environment
python -m venv venv

# Activate virtual environment
# On Windows:
venv\Scripts\activate
# On macOS/Linux:
source venv/bin/activate

# Install dependencies (choose appropriate file under requirements/portable for your hardware)
pip install -r requirements/portable/requirements.txt --upgrade

# Launch server (basic command)
python server.py --portable --api --auto-launch

# When done working, deactivate
deactivate
```

That path is called fast setup on any Python 3.9+, and it ends by launching a server, so the interface arrives in a browser rather than as a window. The one-click route is the alternative: clone or download the source archive, run `start_windows.bat`, `start_linux.sh`, or `start_macos.sh`, pick your GPU vendor when prompted, then open `http://127.0.0.1:7860`. Those scripts use Miniforge to create a Conda environment inside `installer_files/`, and `cmd_linux.sh`, `cmd_windows.bat`, or `cmd_macos.sh` open an interactive shell inside it. The tree also holds a Colab-TextGen-GPU notebook and a `docker/` directory, and the project does not document setting either up, so they arrive as files rather than as supported paths.

## Switching backends without restarting does not mean any model will load

Five inference backends are named: llama.cpp, ik_llama.cpp, Transformers, ExLlamaV3, and TensorRT-LLM, and the interface can switch between backends and models without restarting. That is the easy half. The other half is that each backend wants a different model layout, and the layout is fixed by the build you installed. llama.cpp and ik_llama.cpp read GGUF, the one format the portable build covers. Transformers and EXL3 want the multi-file subfolder arrangement, which needs the full installation. The switcher does not reconcile that and does not warn you when you pick a combination that cannot work. Two of the five named backends are absent from the portable build's hardware list entirely, which reads CUDA, Vulkan, ROCm, and CPU-only. So the no-restart promise holds for a model that is already loaded, while the set of models you can load at all was decided by an installer run earlier. Read the backend list as a capability list for the full install, and check the format of your weights against it before you point the switcher at them.

## Flags are persisted in user_data/CMD_FLAGS.txt, and there are two ways to pass them

Command-line flags reach the application through one of two conventions. You can pass them directly, for example `./start_linux.sh --help`, or persist them in `user_data/CMD_FLAGS.txt`, with `--api` given as the flag that enables the API. The virtual environment route takes its flags on the command line too, in `python server.py --portable --api --auto-launch`. Two launchers, two conventions, and one of the two is a file that survives restarts. That is convenient until the file becomes the thing nobody remembers editing. It is maintained by hand, it applies on every run of that script, and a flag added to turn the API on therefore stays on until somebody removes the line. The project publishes no list of which flags are valid, and its own example for finding out is asking a script for help. Nothing in the setup instructions says what an unrecognised flag does, whether it is ignored or fatal, so the safe way to find out is to ask a script rather than to guess in the file that governs every later launch.

## Extension requirements land in the same environment, and conflicts are settled by overwriting the app

Extensions are Python code running in the same interpreter, and the project is open about how their requirements are handled. A tool-calling tool is a single `.py` file each, and the built-in and community extensions cover TTS, voice input, and translation from a separate extensions directory repository. To install what they need, you use the update wizard's Install/update extensions requirements option, and the stated reason it ends the way it does is that it reinstalls the main project requirements at the end to ensure they take precedence over conflicting extension dependencies. Read plainly, that is the conflict policy: the application's own requirements win, and the way there is to overwrite the environment the application is running in. The same option is reachable in the automated path through the `INSTALL_EXTENSIONS` environment variable, alongside `GPU_CHOICE` and `LAUNCH_AFTER_INSTALL`, with the example `GPU_CHOICE=A LAUNCH_AFTER_INSTALL=FALSE INSTALL_EXTENSIONS=TRUE ./start_linux.sh`. What the project does not offer is isolation, pinning, or rollback between an extension's dependencies and the application's, so a conflicting extension changes the ground under the app rather than sitting beside it.

## Zero telemetry sits next to a built-in web search tool and a server API

The privacy claim and the network features are both stated plainly, and they describe different things. The project says the application is 100% offline and private, with zero telemetry, external resources, or remote update requests. It also ships tool-calling in which models call custom functions during chat, including web search, page fetching, and math, with MCP servers supported, and it exposes an OpenAI/Anthropic-compatible API with Chat, Completions, and Messages endpoints for other programs to call. The first statement is about what the application sends on its own initiative. The second is about what it can reach once you hand it a tool or point another client at it. The gap for a reader is that switching on a web search tool changes who is responsible for the request, and the project does not document what that tool sends, to which endpoint, or how to point it at a local index instead. Offline describes the default configuration rather than a property you keep after enabling a tool.

## Modes select a Jinja2 template, not a sampling parameter

Modes in this application choose a prompt template rather than a tuning knob. There is instruct mode for instruction-following, described as behaving like ChatGPT, and chat-instruct and chat modes for talking to custom characters, and the project states that prompts are automatically formatted with Jinja2 templates. That one fact explains most of what people try to configure and cannot. The template comes from the model you loaded, the mode tells the formatter which one to apply, and when the two do not match, the model receives a layout it was never trained on and there is no setting that repairs it. The project does not document overriding the template per model, so the way to change a prompt's shape is to switch modes or change models. The same formatting pass covers file attachments, which take text files, PDF documents, and `.docx` documents, and the message controls, which allow editing a message, moving between message versions, and branching the conversation at any point. The Notebook tab is the escape hatch, sitting outside chat turns for free-form text generation.

## Conclusion

TextGen suits a developer who wants a local interface, a switchable backend, and an OpenAI or Anthropic compatible endpoint on their own machine, and the portable folder is the cheapest way to find out. It fits badly if your models are multi-file Transformers or EXL3, because that route requires the full installation and its PyTorch download, and it fits badly if you need extension isolation, since the documented remedy for a dependency conflict is to reinstall the application's own requirements. Before committing, check which model format each backend you plan to use wants, look at what your flags in user_data/CMD_FLAGS.txt are actually enabling, and decide what any web search tool you switch on is allowed to reach.

## FAQ

### How do I use TextGen?

Download, unzip, and double-click textgen, then download a GGUF model file from Hugging Face and place it in the user_data/models folder, where the interface detects it automatically. The script route is start_windows.bat, start_linux.sh, or start_macos.sh, after which you open http://127.0.0.1:7860 in a browser.

### What is a text generation app?

In this project's terms it is a desktop app for local LLMs with instruct mode for instruction following, chat-instruct and chat modes for custom characters, a notebook tab for free-form text outside chat turns, and vision support that attaches images to messages for visual understanding.

### Which text generator is considered the best?

The project does not rank text generators. It names the backends it can run, llama.cpp, ik_llama.cpp, Transformers, ExLlamaV3, and TensorRT-LLM, and says you can switch between backends and models without restarting, which leaves the choice to be made per model format and per installed build.

### How does TextGen webui compare with Ollama?

The repository does not make that comparison. What it states about itself is that it is 100% offline and private with zero telemetry, external resources, or remote update requests, runs five named backends, and ships an OpenAI/Anthropic-compatible API with Chat, Completions, and Messages endpoints and tool-calling support.

### What is text generation in TextGen?

The project applies it in two places. Inside chat turns, instruct mode handles instruction following while chat-instruct and chat modes handle custom characters, and prompts are formatted automatically with Jinja2 templates. Outside chat, a separate notebook tab exists for free-form text generation.

## Sources

- [Official README](https://github.com/oobabooga/textgen#readme)
- [Project repository](https://github.com/oobabooga/textgen)
- [Release notes](https://github.com/oobabooga/textgen/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/oobabooga-textgen
