Wuffs: A Memory-Safe Language and C Library for Parsing Untrusted File Formats
Wrangling Untrusted File Formats Safely
At a glance
- What is it?
- Wuffs is a memory-safe programming language designed specifically for writing file format parsers, decoders, and encoders. It achieves safety by requiring compile-time proofs of bounds, overflow, and null pointer correctness rather than relying on runtime checks, and it transpiles to a single C file that any C or C++ project can embed without a Go toolchain.
- Who is it for?
- Wuffs is the right choice when you are writing a parser or decoder for an untrusted file format in a C/C++ codebase that cannot tolerate memory-safety bugs, and when you are willing to invest the additional time required to annotate proofs in the source. The performance and safety properties are real: they come from the design of the language itself rather than from tooling bolted on afterward.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 14 days ago.
- What is it written in?
- Mainly C, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Wuffs Solves for C/C++ Developers Handling Untrusted Data
File format parsers are among the most exploited components in software. A malformed PNG, GIF, or ZIP file fed to a hand-written C parser can trigger a buffer overflow, an integer overflow, or a null pointer dereference. These three bug classes are the root cause of a substantial fraction of memory-safety CVEs in C/C++ projects.
Wuffs addresses this by making those three bug classes impossible to express. The compiler rejects any statement that it cannot prove safe. An integer increment is a compile-time error unless the compiler can prove the variable is not already at its maximum value for that type. A buffer access is a compile-time error unless the compiler can prove the index is within bounds. If the code compiles, those three classes of bugs are absent.
The trade-off is explicit: programmers must annotate their code with the proofs the compiler needs. The README describes this directly: 'the trade-off in aiming for both safety and speed is that Wuffs programs take longer for a programmer to write, as they have to explicitly annotate their programs with proofs of safety.'
Wuffs is not a systems programming language for writing applications. It is for writing libraries, specifically the parsing, decoding, and encoding kernels that handle untrusted binary data. The README is clear that the idea is to write only the parts of a larger project that are both performance-conscious and security-conscious in Wuffs, not the entire application.
How Compile-Time Bounds Checking Works in Practice
The Wuffs compiler enforces safety proofs at every step. A concrete example from the README shows what happens when a proof fails. In the LZW decoder for GIF, a one-line change that would cause an integer to potentially overflow produces a compile-time error:
$ wuffs gen std/gif
check: expression "(c + 1) as u8" bounds [1 ..= 256] is not within bounds [0 ..= 255] at
/home/n/go/src/github.com/google/wuffs/std/lzw/decode_lzw.wuffs:101. Facts:
n_bits < 8
c < 256
this.stack[s] == (c as u8)
use_save_codeThe error message names the exact expression, the bounds it would produce, the bounds the type allows, and the facts the compiler used to reach that conclusion. The compiler tracks variable ranges statically and reports exactly what constraint is violated and why.
In contrast, adding a conditional that gates the potentially unsafe operation so the compiler can prove it only executes when safe allows the code to compile, even if the logic introduced by the condition is incorrect. The safety guarantee from Wuffs is narrow: if it compiles, those three specific bug classes (buffer overflows, integer overflows, null pointer dereferences) are absent. Logic errors such as an incorrect condition are not caught.
The compiled output is always C code. Running `wuffs gen std/gif` when the safety proofs pass produces a C file in the `gen/c/` directory. The developer then uses that C file, not the Wuffs source, in downstream projects.
Wuffs the Language Versus Wuffs the Library
The README distinguishes two ways to work with Wuffs. Wuffs the Language is the language and its toolchain, written in Go. Using it means writing `.wuffs` source files, running `wuffs gen` to produce C code, and running `wuffs test` to verify correctness. This workflow requires a Go installation and the Wuffs Go tools.
Wuffs the Library is the output of that workflow: the C code that already exists in the `release/c/` directory of the repository. Any C or C++ project can include that C file directly, without installing Go, without understanding the Wuffs language, and without touching any `.wuffs` source. From the user's perspective, it is just a C library.
The README states this distinction directly: 'Other C/C++ projects can use that library without requiring the Wuffs the Language toolchain. Those projects can use Wuffs the Library like using any other third party C library. It's just not hand-written C.'
For most C/C++ projects that want to add a fast, safe decoder for GIF, PNG, or another supported format, the library form is the appropriate starting point. The language form is for researchers or developers who want to add new format support or modify existing decoders.
Repository Layout and the Example Programs
The repository organizes its source according to the distinction between the language implementation and the library output:
The `lang/` directory holds the Go libraries that implement the Wuffs language: tokenizer, AST, parser, renderer, and related tools. The `lib/` directory holds other Go libraries not specific to the language. The `cmd/` directory holds the command-line tools, including `wuffs gen` and `wuffs test`.
The `std/` directory holds the standard library written in Wuffs itself, including decoders for formats such as GIF, PNG, BMP, JPEG, and compression algorithms such as deflate, bzip2, and LZW. The `release/c/` directory holds the transpiled C output that end users consume.
The `example/` directory contains working example programs that use Wuffs the Library. These include programs for converting between CBOR and JSON, computing CRC32 checksums, playing GIF animations, viewing images, and decompressing files. Each example is a self-contained C program that demonstrates how to link against the library.
The `fuzz/` directory contains fuzzing harnesses. Because Wuffs is designed for untrusted input, the repository ships fuzzing infrastructure, and the go.sum file reflects the test dependencies.
Hermeticity: No Syscalls, No Memory Allocation
Wuffs code is hermetic by design. It cannot make any system calls, which means it cannot read files, open network connections, or interact with the operating system. It cannot allocate or free memory, which eliminates memory leaks, use-after-frees, and double-frees.
The README describes the model as Sans I/O style: the library receives bytes as input and produces bytes as output. The caller is responsible for actual I/O. A Wuffs-based GIF decoder does not open a file; the caller opens the file and passes the bytes to the decoder. A Wuffs-based deflate decompressor does not read from a socket; the caller reads from the socket and passes the compressed bytes to the decompressor.
This design means Wuffs code is agnostic to synchronous versus asynchronous I/O. The library caller chooses how to handle I/O; the library itself does no I/O. The README uses the term 'function color' to describe the synchronous versus asynchronous distinction, and notes that Wuffs code avoids the problem entirely.
The hermeticity constraint is also why Wuffs is explicitly not a general-purpose language. Writing a program, as opposed to a library, requires syscalls for reading input, writing output, and managing process state. Wuffs cannot do those things. The README notes directly: 'it is unlikely that a Wuffs compiler would be worth writing entirely in Wuffs.'
Performance Benchmarks Documented in the Repository
The README documents specific performance comparisons against established C libraries for several formats. These comparisons are part of the repository's documentation, not external claims.
For GIF decoding, the README states Wuffs decodes 2x to 6x faster than giflib (C), image/gif (Go), and gif (Rust). For PNG decoding, the benchmark page linked from the README shows 1.2x to 2.7x faster than libpng (C), image/png (Go), and png (Rust). For deflate, the figure is up to 1.4x faster than zlib. For bzip2, 1.3x faster than /usr/bin/bzcat.
These gains come from a combination of factors: the code is generated C without the overhead of a runtime-checked safe language, the algorithms are written with explicit bounds knowledge that allows the compiler to eliminate checks, and the output is deterministic because there is no garbage collector.
Using the C library form from release/c/ allows a project to benefit from these performance characteristics without writing any Wuffs code. The library is a drop-in addition to a C or C++ project.
Limitations and When Not to Use Wuffs
Wuffs is not a suitable choice for general application development. The hermeticity constraint rules out writing servers, command-line tools, or any software that needs to interact with the operating system directly. Even within the domain of file format handling, Wuffs requires manual proof annotations, which makes development slower than writing equivalent C code that relies on sanitizers for testing.
The language does not yet have wide IDE or editor support. The toolchain is written in Go, and Go 1.25 is required per the go.mod file. Developers who want to write new Wuffs decoders must install Go and understand the Wuffs language semantics well enough to write the required proof annotations.
Wuffs has no GitHub releases, and the most recent push was on 2026-09-16. The lack of formal releases means adopters who use the library form should pin to a specific commit rather than pulling from main, since the C output in release/c/ can change across commits as the standard library evolves.
The format coverage in the standard library is broad for the formats it does cover (GIF, PNG, BMP, JPEG, deflate, bzip2, LZW, JSON, CBOR, and others), but if a project needs a format not in the std/ directory, it either needs to wait for the Wuffs community to add it or write the decoder itself, which requires learning the Wuffs language.
Editorial conclusion
Wuffs is the right choice when you are writing a parser or decoder for an untrusted file format in a C/C++ codebase that cannot tolerate memory-safety bugs, and when you are willing to invest the additional time required to annotate proofs in the source. The performance and safety properties are real: they come from the design of the language itself rather than from tooling bolted on afterward. Wuffs is the wrong choice for general-purpose programming, for writing application-level code that needs to allocate memory or make system calls, or for anyone who needs a quick way to add a decoder without significant implementation effort. Using the library form from release/c/ requires no Wuffs toolchain knowledge and is the recommended path for most C/C++ projects.
Frequently asked questions
What file formats does the Wuffs standard library support?
The README and repository's std/ directory include decoders for image formats (GIF, PNG, BMP, JPEG), compression algorithms (deflate, bzip2, LZW), and data formats (JSON, CBOR). The exact set of supported formats is in the std/ directory of the repository.
Do I need to know the Wuffs language to use the Wuffs C library?
No. The release/c/ directory contains the transpiled C output of the Wuffs standard library. Any C or C++ project can include that file directly without installing the Go toolchain or understanding the Wuffs language.
How does Wuffs prevent buffer overflows at compile time?
The Wuffs compiler tracks variable ranges statically and requires proof annotations at statements that could overflow or exceed bounds. A statement like x += 1 is a compile-time error unless the compiler can prove x is not already at the maximum value for its type. If the code compiles, buffer overflows, integer overflows, and null pointer dereferences in those three classes are absent.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/google-wuffs)