# flex: the scanner generator every C toolchain still depends on

> A scanner generator inherited from Berkeley and Vern Paxson, distributed under a permissive licence, shipped inside most Linux systems whether anyone knows it is there or not.

**westes/flex** — The Fast Lexical Analyzer - scanner generator for lexing in C and C++

- Repository: https://github.com/westes/flex
- Stars: 4,042 · Forks: 568
- Language: C
- License: NOASSERTION
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/westes-flex

## What flex generates and for whom

The README opens with the definition: flex is a tool for generating scanners, programs which recognize lexical patterns in text. The repository description expands it to a scanner generator for lexing in C and C++, with topics for c, flex, lexer and lexer-generator.

The generated artifact is a C source file implementing a table driven state machine with the pattern set compiled in. You describe tokens in a specification file, run flex, and get a `.c` file to compile into your program. Because the output is plain C rather than a runtime library with a data file, the result links into anything with a C compiler, which is the main reason flex has survived while several contemporaries have not.

The `lib/` directory in the tree corresponds to `libfl`, the small support library a generated scanner links against, and the release notes for 2.6.3 explain that compatibility functions in `lib/` are no longer built as a library but as objects, a change made specifically to be friendlier to cross compilation. You can also build flex without building libfl at all, via a `--disable-libfl` configure option.

## Two licences, and which one actually applies

The licensing here needs reading carefully, because it is permissive and the repository's own metadata does not say so.

GitHub reports the license field as NOASSERTION, meaning nothing was detected automatically. The README then states that the code is derived from software contributed to Berkeley by Vern Paxson, and that the United States Government holds rights in the work under contract no. DE-AC03-76SF00098 between the Department of Energy and the University of California.

The licence terms quoted in the README are the three-clause BSD pattern: retain the copyright notice, conditions and disclaimer in source redistributions; reproduce them in the documentation or other materials of binary redistributions; and neither the name of the University nor the names of its contributors may be used to endorse or promote derived products without prior written permission. The warranty disclaimer is the standard as-is language with the implied warranties of merchantability and fitness disclaimed.

The `COPYING` file in the tree is the authoritative copy and is described in the README's file list as flex's copyright and license. A `LICENSE` file is absent from the tree, which explains why automated detection found nothing. If you are vetting flex for a product, `COPYING` is the file to read, and it is a far easier answer than the metadata field suggests.

## Building with autotools, and why autogen.sh exists

The tree is a textbook GNU autotools project: `configure.ac` and `control.ac` at the root, `Makefile.am`, an `m4/` directory for macros, `po/` for internationalization, `autogen.sh` for bootstrapping, and `libt/` style helper tooling under `tools/`. There is also `.indent.pro`, which the 2.6.4 release notes specifically mention when noting that the indent target now knows about flex's newer layout.

The presence of `autogen.sh` means the repository does not ship a generated `configure` script, so you bootstrap before configuring. The 2.6.3 notes describe why that path is friendlier than it used to be: a new `--disable-bootstrap` option changes the build so the scanner is simply built by processing the scanner source, rather than building flex and then using that flex to build flex again. For cross compilation, that default two stage build is awkward, which is what the option is for.

The build's portability work is ongoing rather than historical. Version 2.6.3 fixed a buffer overflow that could occur when the path to m4 was sufficiently long, and the fix removed the dependence on the constant PATH_MAX. Version 2.6.4 records that flex can now be cross compiled, that the configure script has a better idea of which headers are required and will error when missing functions are detected, and that the automake and gettext versions listed as required were lowered. `INSTALL.md` holds basic installation information and `NEWS` carries the current version number plus user visible changes.

## Documentation, examples and the regression suite

The README does something helpful here: it indexes the distribution contents so you know what is in the tarball before you read it.

`doc/` is user documentation and is the manual, referenced from the release notes as containing some typos that were removed in 2.6.4. `examples/` contains examples of some possible flex scanners and a few other things, with its own `examples/README` giving details, and the tree shows the shape of it: `testxxLexer.l`, an `examples/Makefile.am`, a `debflex.awk` helper, a `manual/` directory, and a `fastwc/` directory that looks like a small worked program.

`tests/` is the regression suite, documented in `tests/README`, and the 2.6.2 notes give a sense of its nature: input filenames on MSWindows are calculated correctly, test suite code was cleaned up so it compiles more cleanly, and a new `sv` translation was added under `po/`.

One repository entry is worth knowing about if you are trying to work out what is current: `.prev-version`. Its presence suggests the tree records the previously current version somewhere, alongside `NEWS` and `ONEWS` as two separate news files.

```bash
src/
lib/
doc/
examples/
tests/
```

That listing is the shortest useful answer to what the five top level directories are for, and it is the layout every flex source reader starts from.

## Releases stopped in 2017, the branch did not

The release history is the most surprising thing about this project, and anyone evaluating it should see the dates.

The newest tag is v2.6.4, published on 2017-05-06. Before it, v2.6.3 on 2016-12-30 and v2.6.2 on 2016-10-25. So there have been no published releases in roughly nine years, while the repository's last push was on 2026-08-03 and the build status badge points at an active GitHub Actions workflow.

Both facts are consistent, and the explanation is visible in the release notes themselves. The 2.6.3 entry is almost entirely bug fixes, cross compilation improvements and build system changes; 2.6.2 is a similar mix of internal fixes, test suite corrections and build portability work. These are the changes of a mature tool that has converged.

The practical consequence is a version choice. If you take a packaged flex from your distribution you get whatever that distribution froze. If you build from this repository you get the master branch, which includes everything since 2.6.4 without a version number to distinguish it. `NEWS` is the file that records the current version and the list of user visible changes, so that is where to look rather than the tags list.

## Where to ask questions, and what the mailing lists are for

The README routes support to three places, and distinguishes them by audience rather than sending everyone to one inbox.

Bugs and patches go through GitHub's issues and pull request features. Usage questions go to `flex-help@lists.sourceforge.net`, development discussion to `flex-devel@lists.sourceforge.net`, and release announcements to `flex-announce@lists.sourceforge.net`.

The SourceForge page at sourceforge.net/p/flex/mailman/ is where you subscribe or search the archive, and the README notes that posting is only allowed from subscribed addresses. That constraint is the one thing that can stop a question from landing, so it is worth reading before composing anything.

For documentation questions specifically, the local files come first. `doc/` holds the manual, `examples/README` describes the sample scanners, `tests/README` explains the regression suite, and `ABOUT-NLS` describes the internationalization support that the `po/` directory implements. A build problem is most efficiently answered from `INSTALL.md`, and a contribution from `CONTRIBUTING.md`.

## Conclusion

flex is infrastructure rather than a project you choose, and that shapes both its virtues and its caveats. It has been stable for so long that the generated scanners compile with almost any C toolchain, it has no dependencies beyond a C compiler and the autotools it ships with, and its permissive licence lets you build it into commercial software without the copyleft obligations that would make that awkward. The caveat is the same one the release dates suggest: the newest tag is 2.6.4 from 2017, so bug fixes reach the master branch without a release. Check the `NEWS` file for the current version number and build from a checkout rather than expecting a new tarball. The documentation is in `doc/`, worked examples are in `examples/` with its own README, and `tests/` holds the regression suite. If you need the flex that ships with your distribution, your package manager already has it, and building it yourself is worthwhile mainly when you need a newer fix.

## FAQ

### What is flex (lexical analyzer generator)?

flex is a tool for generating scanners, which are programs that recognize lexical patterns in text. You write patterns in a specification file, run flex, and it produces a C source file with the pattern set compiled into a table driven state machine. The repository describes it as a scanner generator for lexing in C and C++, and the generated output is plain C with a small support library called libfl.

### What are the key differences between lex and flex?

flex descends directly from lex, deriving from software contributed to Berkeley by Vern Paxson, so the specification language and the generated scanner structure are close relatives rather than a redesign. The differences are in implementation and tooling: flex is maintained as an autotools project with GNU extensions, its generated scanners use dynamic tables rather than the old one-array-per-state scheme, and it adds features such as interactive and yy_scan-based input and reentrant scanners. In practice most lex specifications move to flex with minimal edits.

### What is the purpose of a lexical analyzer?

A lexical analyzer turns a stream of characters into a stream of tokens, so the parser above it never deals with raw text. It recognizes patterns such as identifiers, numbers, string literals and comments, and reports each match with its type and value. That is exactly the job the README describes for the scanners flex generates, which is why lexing is the first stage after reading input in most compilers and text tools.

## Sources

- [Issues](https://github.com/westes/flex/issues)
- [README](https://github.com/westes/flex/blob/master/README.md)
- [Releases](https://github.com/westes/flex/releases)
- [westes/flex on GitHub](https://github.com/westes/flex)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/westes-flex
