# Simple Transformers: a wrapper for training HuggingFace models in three calls

> Simple Transformers wraps the HuggingFace Transformers library behind task-specific model classes. It is aimed at people who want to fine-tune a transformer without writing a training loop, and the price is a layer of abstraction over a fast-moving dependency.

**ThilinaRajapakse/simpletransformers** — Transformers for Information Retrieval, Text Classification, NER, QA, Language Modelling, Language Generation, T5, Multi-Modal, and Conversational AI

- Repository: https://github.com/ThilinaRajapakse/simpletransformers
- Website: https://simpletransformers.ai/
- Stars: 4,256 · Forks: 710
- Language: Python
- License: Apache-2.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/thilinarajapakse-simpletransformers

## The gap Simple Transformers fills between Transformers and a training script

HuggingFace Transformers gives you model classes, tokenizers and trainers. What it does not give you is a task-shaped object that already knows which columns to read from a DataFrame, which metric to compute at the end, and how to write predictions back out. Simple Transformers is that object. The README states the pitch directly: "Only 3 lines of code are needed to initialize, train, and evaluate a model." The library is built on top of Transformers rather than replacing it, so the underlying weights, tokenizers and model hubs are the same ones you would use by hand.

The audience is narrow and identifiable. It is not a library for researchers who want to modify attention or write a custom scheduler. It is for engineers and applied scientists who have a labelled dataset, a deadline, and no interest in reimplementing the loop that feeds batches to a model. The task list is the real scope statement: information retrieval (dense retrieval), language model training and generation, encoder fine-tuning, sequence classification, token classification (NER), question answering, language generation, T5 and other seq2seq tasks, multi-modal classification, and conversational AI. If your problem is not on that list, the wrapper has nothing for you.

One design consequence is worth naming early. Because each task gets its own class, the API is not uniform across tasks. The README says the high-level pattern is the same everywhere, but that "there are necessary differences between the different models" in input and output data formats and in task-specific configuration options. So the mental model is not one API with a task parameter. It is a family of classes that share method names.

## How the task-specific model classes work

Every wrapped model follows four steps, in this order: initialize a task-specific model, train with train_model(), evaluate with eval_model(), and predict with predict(). The class you pick determines the data shape. ClassificationModel handles binary and multi-class text classification, regression, and sentence-pair classification. MultiLabelClassificationModel handles multi-label. NERModel handles token classification. QuestionAnsweringModel, LanguageModelingModel, LanguageGenerationModel, ConvAIModel, MultiModalClassificationModel, RepresentationModel and RetrievalModel cover the rest.

Data flows in as a pandas DataFrame. In the README example the training frame has exactly two columns, text and labels, and the labels are integers. That column contract is the part people get wrong first: the wrapper does not infer your schema, it expects the one documented for the task. Configuration flows in through an args object, for example ClassificationArgs, which carries training settings such as num_train_epochs. Model selection flows in through two strings at construction time, a model type and a model name, as in the roberta and roberta-base pair in the README.

Evaluation returns a tuple. In the README example that is result, model_outputs, wrong_predictions. The third element is the useful one in practice: it gives you the individual examples the model got wrong, which is where you look when a metric is lower than expected. Predictions come back as a tuple too, predictions and raw_outputs, so you get both the decoded answer and the underlying scores.

Underneath, this is a thin layer. setup.py declares transformers>=4.31.0 as a dependency, along with datasets, tokenizers, scikit-learn, seqeval, pandas, tensorboard, tensorboardx, wandb, streamlit and sentencepiece. That dependency list is what the wrapper is actually orchestrating.

## Installing Simple Transformers and running a first classification job

The README gives a conda-first setup. Create an environment with python, pandas and tqdm, then add PyTorch. The CUDA variant pins a specific toolkit version, and there is a CPU-only command for machines without a GPU.

```bash
conda create -n st python pandas tqdm
conda activate st
conda install pytorch>=1.6 cudatoolkit=11.0 -c pytorch
pip install simpletransformers
```

If you do not need CUDA, the README substitutes conda install pytorch cpuonly -c pytorch for the third line. Either way, pip install simpletransformers is the step that pulls the wrapper itself. The repository also has a Makefile target, install, which runs pip install -e . followed by pip install -r requirements-dev.txt; that path is for working on the library, not for using it.

Once installed, the README example is a complete runnable script. It builds two tiny frames with text and labels columns, sets one argument, constructs the model, and calls the three methods.

```python
from simpletransformers.classification import ClassificationModel, ClassificationArgs
import pandas as pd

train_df = pd.DataFrame([
    ["Aragorn was the heir of Isildur", 1],
    ["Frodo was the heir of Isildur", 0],
], columns=["text", "labels"])
eval_df = pd.DataFrame([
    ["Theoden was the king of Rohan", 1],
    ["Merry was the king of Rohan", 0],
], columns=["text", "labels"])

model_args = ClassificationArgs(num_train_epochs=1)
model = ClassificationModel("roberta", "roberta-base", args=model_args)
model.train_model(train_df)
result, model_outputs, wrong_predictions = model.eval_model(eval_df)
```

What you should see: the model downloads roberta-base weights on first run, trains for one epoch over four examples, and eval_model returns a metrics dictionary plus the per-example outputs and the misclassified rows. The README's snippet ends mid-call at model.pre, so the prediction line is not fully shown there; the documented method name is predict().

For tracking, the README lists an optional step: pip install wandb, and the setup.py dependency list already includes wandb>=0.10.32. The README does not document what happens if wandb is installed but not configured, so treat that as unverified.

## Where the wrapper gets in your way

The abstraction is the limitation. Anything the wrapper does not surface as an argument or a method is either reachable through the underlying Transformers objects or not reachable at all, and the documentation does not promise which. If you need a custom loss, a bespoke sampling strategy, or a training loop that interleaves two models, you are fighting the layer rather than using it.

The second constraint is version coupling. setup.py requires transformers>=4.31.0 with no upper bound. Transformers changes its internals regularly, and a wrapper that calls into those internals inherits the breakage. The README points to the Changelog for changes, which is the right place to look, but it also means upgrading Transformers is a decision you make deliberately rather than something you let pip resolve.

The third is the release cadence visible in the repository. The most recent release listed is v0.60.0 (New Classification Models) from 2021-02-01, while setup.py declares version 0.70.8 and the last push to the default branch was on 2026-05-31. That gap between tagged releases and the working tree is worth understanding before you pin anything: if you install from PyPI you get a published version, and if you install from the repository you get something newer than the release notes describe.

Finally, the README's own example is four sentences long. That is a demonstration of the API, not of model quality. Nothing available here indicates how these models perform on real datasets, and no benchmark numbers are given.

## Simple Transformers against writing the Transformers loop yourself

The honest alternative is not another wrapper. It is using HuggingFace Transformers directly, with its Trainer class or a hand-written loop. The difference is where the work sits. With Transformers alone you write the dataset class, the tokenization call, the metric function and the training arguments. In exchange you control every one of those pieces and you are coupled to one dependency instead of two.

Simple Transformers moves that work into the library. You get the four-step pattern and a per-task class, and you accept the library's choices about data format, metrics and defaults. For a classification or NER baseline this is a good trade. For anything where the metric is the research contribution, it is usually not, because you will end up overriding the metric anyway.

There is a middle position worth noting from the repository layout. The examples directory is organized by task, with separate folders for text_classification, named_entity_recognition, question_answering, language_generation, language_representation, llms, retrieval, seq2seq, t5 and hyperparameter tuning. Those scripts are the practical documentation for anything the README does not cover, and they sit closer to the code than the prose does. The project also ships a bin/simple-viewer script, registered in setup.py, which is not described in the README.

## Maintenance, upgrades and the Apache-2.0 licence

The repository is not archived, and the last push to the default branch was on 2026-05-31. That is the only maintenance signal available here, and it should be read alongside the release history rather than instead of it: the newest release named is from 2021, while the package version in setup.py is 0.70.8. Anyone pinning a version should decide whether they are pinning the PyPI release or the repository state, because they are not the same thing.

Upgrade cost has two parts. The first is Transformers itself, which is declared as transformers>=4.31.0. A major Transformers release can change the internals this wrapper depends on. The second is the rest of the dependency set: datasets, tokenizers, scikit-learn, seqeval, tensorboard, tensorboardx, wandb, streamlit and sentencepiece. Several of those are heavy, and streamlit in particular is pulled in as a hard dependency even if you never open a browser UI. A minimal deployment that only needs classification still installs the full list.

On licensing: the project is Apache-2.0, and the LICENSE file is at the repository root. That is a permissive licence with an explicit patent grant, which is generally the easy case for commercial use, but it governs this library only. The pretrained weights you load through it, such as roberta-base, carry their own terms from whoever published them, and nothing in this repository changes those. That is a question for your own legal review, not something the README answers.

## Conclusion

Adopt Simple Transformers when your task is one of the wrapped ones (classification, NER, QA, language modelling, generation, retrieval) and you want a working baseline today rather than a custom training loop. Do not adopt it if you need a model architecture or training schedule the wrapper does not expose, or if you cannot accept the version coupling to transformers>=4.31.0. Before committing, run the example script for your task from the examples directory and check that eval_model() returns the metric you intend to report.

## FAQ

### What is Simple Transformers?

It is a Python library that wraps HuggingFace Transformers with task-specific model classes, so you can initialize, train and evaluate a model in three calls. It supports classification, NER, question answering, language modelling, generation, retrieval and multi-modal tasks.

### Does ChatGPT use transformers?

The README does not discuss ChatGPT or any OpenAI model. It only states that Simple Transformers is based on the HuggingFace Transformers library and that it wraps that library's models.

### Can you explain transformers in a simple way?

Simple Transformers does not explain the transformer architecture. It treats the model as something you construct from a model type and a model name, such as roberta and roberta-base, and then train and evaluate.

### What are the three types of transformers?

The README does not divide transformers into three types. It lists the tasks the library supports, including information retrieval, sequence classification, token classification (NER), question answering, language generation, seq2seq and multi-modal classification.

## Sources

- [License: Apache-2.0](https://github.com/ThilinaRajapakse/simpletransformers/blob/master/LICENSE)
- [Project website](https://simpletransformers.ai/)
- [README](https://github.com/ThilinaRajapakse/simpletransformers/blob/master/README.md)
- [Releases](https://github.com/ThilinaRajapakse/simpletransformers/releases)
- [ThilinaRajapakse/simpletransformers on GitHub](https://github.com/ThilinaRajapakse/simpletransformers)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/thilinarajapakse-simpletransformers
