TextAttack: named attack recipes, one cache directory, and a docs-check target that duplicates docs
TextAttack 🐙 is a Python framework for adversarial attacks, data augmentation, and model training in NLP https://textattack.readthedocs.io/en/master/
At a glance
- What is it?
- A Python framework that turns adversarial attacks, data augmentation and model training into command line work. An attack is described as a goal function, a set of constraints, a transformation and a search method, and the published attacks from the literature ship as named recipes you can list. Everything downloaded lands in one cache directory you can move, and the repository has not published a release since a two year gap.
- Who is it for?
- TextAttack suits a research or applied NLP team that wants to reproduce published attacks against its own models, or to use adversarial examples as augmentation, without writing the search loop by hand. It does not suit a production service that needs a small dependency footprint, because the base requirements pull in a sequence labelling library and a set of linguistic tooling packages, and several of the extras are pinned to old versions.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 51 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Four jobs, one command, and two entry points
The README gives four reasons to use the framework and they are four different jobs rather than one: understanding models better by running attacks against them and reading the output, researching and developing new attacks with a library of components, augmenting a dataset to improve generalisation and robustness downstream, and training a model with a single command including all downloads. The interface is one command with subcommands, of which two are called common, one to run attacks and one to augment. Every command takes a help flag, and the whole tool can be invoked either as a console script or as a Python module:
pip install textattackThe examples directory is split four ways to match those jobs, with separate folders for attacks, augmentation, datasets and training, including a script for augmenting a spreadsheet file.
A recipe is four components, and the table is the documentation
The most useful page in the README is the recipe table, because it defines each published attack as a tuple of four things rather than as a name you have to look up. Each row names the goal function, the constraints enforced, the transformation and the search method. Reading the first rows makes the structure obvious: one attack combines untargeted classification with constraints on the percentage of words perturbed, word embedding distance, sentence encoding similarity and part of speech consistency, applies either a counter-fitted embedding swap or masked token prediction, and searches greedily by gradient. Another applies the same kind of transformation with a genetic algorithm as the search. Recipes are listed with a dedicated list command and run by name, which means you can enumerate what exists before deciding which one to reproduce.
Everything downloaded lands in one directory you can move
The setup section has a tip that matters more than it looks. TextAttack downloads files to a cache directory in your home folder by default, and that includes pretrained models, dataset samples and the configuration file itself. To change it you set an environment variable, and the example given sets it to a temporary path before running an attack. Two consequences follow. First, a container or a shared machine can be pointed at a volume instead of a home directory, which is the difference between a cache that persists across runs and one that re-downloads a model each time. Second, because the configuration file sits in the same place, anything you change there is as versionless as everything else, so the cache is part of your environment rather than part of your repository.
Two flags change the shape of a run: --parallel and --interactive
Most of the command line surface is one flag at a time and the flags are additive, but two of them change what the tool is doing rather than how it looks. The parallel option distributes an attack across multiple GPUs, and the README is specific that this can really help performance for some attacks, while pointing at a separate example script for attacking Keras models in parallel, which suggests that path needed different handling. The interactive flag replaces the dataset and the example count with samples typed by the user, which turns the tool into a probe you drive by hand. Both are worth knowing before you launch a long run, because discovering either one afterwards means throwing the run away.
The release gap is two years, and the version number lives in the docs config
Three releases are visible in the history and the spacing between them is the story. The most recent is version 0.3.11, published in August 2026, and before it came 0.3.10 in March 2024 and 0.3.9 in September 2023, so there is a gap of more than two years between the middle and the newest tag, with the last commit on the default branch dated 15 August 2026, the day before that release. The packaging explains where the number comes from: the setup file does not carry a literal version at all, it imports the release value from the documentation configuration module, with a comment saying the version is tracked there. So the version of an installed package is defined by a file inside the docs build, which is unusual and worth knowing if you script anything against it.
The base requirements pull in a sequence labeller and Chinese NLP tooling
The dependency list is longer than a targeting library needs, and reading it tells you what the framework does beyond running attacks. There is a sequence labelling library in the base requirements rather than in an extra, along with a grammar checking package, an inflection library, an edit distance package, a semantic similarity score, and two Chinese language packages plus a dictionary resource for Chinese word segmentation and pinyin. The deep learning stack is the expected pair, with a minimum version on the tensor library, an explicit exclusion of one specific early version, and a minimum on the transformer library that is far newer than the Python floor the README states. Optional extras exist for sentence embeddings, a linguistic annotation toolkit, two experiment trackers and topic modelling, and there is a separate extra for the tensor flow stack.
The docs-check target runs the same command as docs
The Makefile is a small tour of the toolchain and contains one bug worth reporting. Formatting runs three tools in sequence, with the import sorter in atomic mode over the test and package directories and a docstring formatter over both. Linting runs the same formatter in check mode, the sorter in check mode, and a style checker with a hard coded list of ignored codes and a directory exclusion. Tests run under pytest with distribution by file and as many workers as the machine has. Then there are the documentation targets: one builds the HTML, one is supposed to build it and exit with an error code rather than a warning, and one runs a live reload server on a fixed port. The problem is that the first two targets contain the identical command, so the stricter one does nothing extra.
Editorial conclusion
TextAttack suits a research or applied NLP team that wants to reproduce published attacks against its own models, or to use adversarial examples as augmentation, without writing the search loop by hand. It does not suit a production service that needs a small dependency footprint, because the base requirements pull in a sequence labelling library and a set of linguistic tooling packages, and several of the extras are pinned to old versions. Three things to check before you build a pipeline on it: which recipe matches the attack in the paper you are reproducing, since the table is the only definition of what each one enforces; whether your hardware will make the run bearable, since a GPU is called optional but said to greatly improve speed and there is a flag for spreading an attack across several; and where the cache goes, since models and dataset samples are downloaded by default into a single directory in your home folder.
Frequently asked questions
What is TextAttack?
TextAttack is a Python framework for adversarial attacks, data augmentation and model training in natural language processing, described more narrowly in its packaging as a library for generating text adversarial examples. It is maintained by a lab at the University of Virginia, released under the MIT licence, and installed with pip.
How do I run an attack with TextAttack?
Through the command line, with the attack subcommand and its help flag. You can either name a recipe that implements a published attack, or spell out the model, search method, transformation, constraints and goal function yourself. A worked example in the README runs a named recipe against a BERT sentiment model over a hundred examples, and passing the interactive flag lets you type your own samples instead of naming a dataset.
What is an attack recipe in TextAttack?
A named combination of four components: a goal function, a set of enforced constraints, a transformation and a search method. The README publishes a table defining every recipe this way, including the constraint on the percentage of words perturbed, embedding distance, sentence similarity or part of speech that each one enforces, and cites the paper each came from. Recipes are listed with a dedicated command and run by name.
Where does TextAttack store downloaded files?
In a cache directory under your home folder by default, which holds pretrained models, dataset samples and the configuration file. Setting the cache directory environment variable relocates all of it, and the README gives an example that points it at a temporary path before running an attack.
How do I use multiple GPUs with TextAttack?
With the parallel option, which distributes an attack across your GPUs and, the README notes, can really help performance for some attacks. For models built with Keras the README points instead to a dedicated parallel example script, which suggests that path is handled separately.
What Python version does TextAttack require?
The README says you should be running Python 3.6 or newer, and that a CUDA compatible GPU is optional but will greatly improve code speed. The formatter in the project configuration targets Python 3.9 through 3.11, so the code style baseline is newer than the stated minimum.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/qdata-textattack)