# xTuring: fine-tuning open-source LLMs with LoRA and INT4

> xTuring is a Python library from Stochastic.ai that wraps data preparation, fine-tuning, evaluation and inference for open models such as LLaMA 2, Mistral, Qwen3 and GPT-OSS. Its value is the short API; its cost is that the published release is from 2023 while the main branch keeps moving.

**stochasticai/xTuring** — Build, personalize and control your own LLMs. From data pre-processing to fine-tuning, xTuring provides an easy way to personalize open-source LLMs. Join our discord community: https://discord.gg/TgHXuSJEk6

- Repository: https://github.com/stochasticai/xTuring
- Website: https://xturing.stochastic.ai
- Stars: 2,675 · Forks: 212
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/stochasticai-xturing

## The gap xTuring fills between a dataset and a usable model

Fine-tuning an open model is not one task. It is dataset shaping, a training loop, adapter configuration, quantization settings, and then a way to call the result. xTuring puts those behind one object. The README describes the library as making it simple, fast and cost-efficient to fine-tune open-source LLMs on your own data, locally or in a private cloud, and the pyproject description is narrower: fine-tuning, evaluation and data generation for LLMs.

The intended reader is a Python developer who already has instruction data and does not want to assemble a training stack. The repository ships an examples directory with datasets, features, models, notebooks and a playground UI, so the project expects people to copy a runnable script rather than read an API reference from end to end. Private deployment is stated as a default: run locally or in your VPC. That matters if the instruction data is the sensitive part.

## How the API, dataset and quantization layers connect

The surface is deliberately small. InstructionDataset reads a folder of Alpaca-format data. BaseModel.create takes a string name and returns a model object. That object exposes finetune, generate and evaluate. The README shows evaluate returning perplexity, and it states that perplexity is the only metric currently supported, so evaluation is a sanity check rather than a model comparison suite.

Model choice is a registry lookup, not a class hierarchy you configure. The README lists off-the-shelf, INT8, LoRA, LoRA+INT8 and LoRA+INT4 variants for LLaMA 2 and GPT-OSS, and separate names such as llama2_int8 and qwen3_0_6b_lora. GenericLoraKbitModel is the escape hatch: it takes a Hugging Face model id directly, which is how the README fine-tunes mistralai/Mistral-7B-Instruct-v0.2 in INT4. The INT4 and INT8 paths depend on bitsandbytes, pinned at 0.41.1 in pyproject.toml, and on the load_in_8bit and load_in_4bit loading kwargs. This is the architecture's main constraint, and the README states it plainly: do not upgrade to transformers 5.x yet, because 5.x drops those kwargs and nothing currently shipped needs 5.x. Qwen3-Omni support, which would require transformers>=5.0.0, is not released.

## Install xTuring and run a first fine-tune

Installation is a single package. The README notes that xTuring requires transformers>=4.36.0, so check that constraint before pinning anything else in the same environment.

```bash
pip install xturing
```

The README's quickstart then loads a toy Alpaca dataset from the examples tree and creates the qwen3_0_6b_lora checkpoint, which it describes as lightweight and CPU-friendly. After finetune, generate takes a list of prompts and returns outputs. Expect the first run to spend most of its time downloading weights. The README points to examples/models/qwen3/qwen3_lora_finetune.py for a runnable version of the same flow.

```python
from xturing.datasets import InstructionDataset
from xturing.models import BaseModel

dataset = InstructionDataset("./examples/models/llama/alpaca_data")
model = BaseModel.create("qwen3_0_6b_lora")
model.finetune(dataset=dataset)
output = model.generate(texts=["Explain quantum computing for beginners."])
print(f"Model output: {output}")
```

Running from source is documented separately. The README installs the package in editable mode with development dependencies and then requires pre-commit hooks, including a commit-msg hook, before contributing. That last step is easy to skip and will affect whether a pull request is accepted.

```bash
git clone https://github.com/stochasticai/xturing.git
cd xturing
pip install -e .
pip install -r requirements-dev.txt
pre-commit install
pre-commit install --hook-type commit-msg
```

## Where xTuring is the wrong tool

The release history is the first warning. The newest published release is v0.1.8 from 2023-09-07, while the repository's last push was on 2026-09-08. A pip install gives you the 2023 wheel, not the GPT-OSS or Qwen3 work described in the README. If you need those checkpoints, you are installing from the repository, and the README does not document a supported upgrade or rollback path for that case.

The second limit is dependency friction. transformers is pinned below 5.x by design, bitsandbytes is pinned at 0.41.1, and datasets is pinned at 2.14.5. Any other library in the same environment that wants newer versions will fight these. The third is scope: perplexity is the only evaluation metric the README lists, so xTuring will not tell you whether a fine-tune is better at a task, only whether the language modelling loss moved. If you need benchmark harnesses, preference data or RLHF, this is not the layer for it. Finally, the project classifies itself as Development Status 3 - Alpha in pyproject.toml, and the README does not claim production support.

## How xTuring differs from writing the PEFT loop yourself

The obvious alternative is the Hugging Face stack directly: transformers for loading, peft for adapters, trl for supervised fine-tuning, bitsandbytes for quantization. That path has more moving parts and no shared model registry, but every component is versioned and documented independently, and you are not waiting on a wrapper to expose a new checkpoint. The difference in approach is where the abstraction sits. xTuring hides the training loop behind finetune and the model choice behind a string; the Hugging Face route hands you the loop and expects you to configure it.

A second alternative is a managed fine-tuning service. xTuring's stated position is the opposite one: private by default, run locally or in your VPC. If your data cannot leave your infrastructure, the managed route is closed to you regardless of convenience. The trade is that you own the GPU, the dependency pins and the debugging.

## Licence, maintenance and what an upgrade costs

The project is Apache-2.0, declared in pyproject.toml and shipped as a LICENSE file, with the classifier OSI Approved :: Apache Software License. That is a permissive licence, and it means the library can be used in commercial settings without a separate agreement. It does not cover the model weights you fine-tune: LLaMA 2, Mistral and GPT-OSS each carry their own terms, and nothing in this repository changes those. Check the model card before you ship anything.

The maintenance picture is mixed and worth stating precisely. The repository is not archived, and its last push was on 2026-09-08, so the main branch is being touched. The published release, however, is v0.1.8 from 2023-09-07. That gap is the upgrade cost: features described in the README are ahead of what pip resolves. The README does not document a versioning policy or a deprecation path for the transformers 5.x transition, so the timing of that move is unknown from the documentation alone. If you adopt xTuring, pin your versions in your own requirements file and treat the repository, not PyPI, as the source of truth for recent model support.

## Conclusion

Adopt xTuring if you want a short Python path from an Alpaca-format instruction file to a LoRA adapter, and you accept pinning transformers below 5.x to keep the INT8 and INT4 loaders working. Skip it if you need a supported tool with a recent release, since the newest published version is 0.1.8 from 2023-09-07 even though the main branch was pushed on 2026-09-08. Before committing, install it in a throwaway environment and check that BaseModel.create resolves the checkpoint name you intend to train, because the registry is where model support is decided.

## FAQ

### How do I install xTuring?

The README gives a single command, pip install xturing, and notes that xTuring requires transformers>=4.36.0. For contributing or running from source, the README instead clones the repository, installs it with pip install -e . plus requirements-dev.txt, and runs pre-commit install.

### Which models can xTuring fine-tune?

The README lists GPT-OSS, LLaMA and LLaMA 2, Qwen3, MiniMax M2, GPT-J, GPT-2, DistilGPT-2 and Mamba, with off-the-shelf, INT8, LoRA, LoRA+INT8 and LoRA+INT4 configurations. GenericLoraKbitModel accepts a Hugging Face model id directly, which the README uses with mistralai/Mistral-7B-Instruct-v0.2. Qwen3-Omni support is not released yet.

### Why can I not upgrade transformers to 5.x with xTuring?

The README states that transformers 5.x drops the load_in_8bit and load_in_4bit loading kwargs that the INT8 and INT4 engines rely on, and that nothing currently shipped needs 5.x. Qwen3-Omni, which would require transformers>=5.0.0, is listed as unreleased.

## Sources

- [License: Apache-2.0](https://github.com/stochasticai/xTuring/blob/main/LICENSE)
- [Project website](https://xturing.stochastic.ai)
- [README](https://github.com/stochasticai/xTuring/blob/main/README.md)
- [Releases](https://github.com/stochasticai/xTuring/releases)
- [stochasticai/xTuring on GitHub](https://github.com/stochasticai/xTuring)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/stochasticai-xturing
