# rga (ripgrep-all): searching inside PDFs, archives and Office files

> rga wraps ripgrep with adapters that extract text from PDFs, docx, epub, sqlite, zip and tar files, so a single regex search reaches inside containers. It is a useful layer, but the extraction depends on external binaries you have to install yourself.

**phiresky/ripgrep-all** — rga: ripgrep, but also search in PDFs, E-Books, Office documents, zip, tar.gz, etc.

- Repository: https://github.com/phiresky/ripgrep-all
- Stars: 9,864 · Forks: 219
- Language: Rust
- License: NOASSERTION
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/phiresky-ripgrep-all

## The gap rga fills: grep stops at the file boundary

ripgrep searches text files fast and recurses through directories, but it treats a PDF, a docx or a zip as opaque bytes. If the string you want sits inside a compressed archive or a binary document format, ripgrep will not find it, and the usual workaround is to extract everything to a scratch directory first. rga exists to remove that step. The README describes it as "a line-oriented search tool that allows you to look for a regex in a multitude of file types" that "wraps the awesome ripgrep and enables it to search in pdf, docx, sqlite, jpg, movie subtitles (mkv, mp4), etc."

The audience is narrow but real: people who keep documentation, research papers, mail archives, sqlite databases or media subtitles on disk and want one regex to cover all of it. A developer grepping a repository does not need rga. Someone searching a directory of downloaded papers and ebooks does.

## How the adapter model works

rga does not parse any of these formats itself. It ships a set of adapters, each of which knows how to turn one input type into plain text or into a stream of further files. The README lists the built-in ones. The pandoc adapter shells out to pandoc with the arguments shown in the README (`pandoc --from= --to=plain --wrap=none --markdown-headings=atx`) and covers .epub, .odt, .docx, .fb2, .ipynb, .html and .htm. The poppler adapter runs `pdftotext - -` for .pdf. The ffmpeg adapter extracts metadata, chapters, subtitles and lyrics from .mkv, .mp4, .avi, .mp3, .ogg, .flac and .webm. The zip and tar adapters read the container as a stream and recurse into its contents, and the decompress adapter handles .gz, .bz2, .xz, .zst and related extensions by running another extractor on the decompressed stream. The sqlite adapter uses sqlite bindings to render a database as plain text.

That recursion is the interesting part. A zip containing a tar.gz containing a PDF is not a special case: the zip adapter yields inner entries, the decompress adapter unwraps the gzip layer, the tar adapter lists members, and the poppler adapter turns the PDF into text. rga also caches extracted text in a database under the platform cache directory so repeated searches over the same files are faster, and `--rga-no-cache` disables that entirely.

Two defaults are worth knowing. Adapter selection is by file extension unless you pass `--rga-accurate`, which detects the mime type from the first 8KiB of magic bytes. The README explains why the limit exists: detection cannot always seek on the input when the file is inside an archive. The mail adapter (.mbox, .mbx, .eml) is disabled by default and must be turned on with `--rga-adapters=+mail`.

## Installing rga and running a first search

On Arch, Gentoo and Nix the package manager handles everything. On Debian-based systems the README says to download the rga binary from GitHub Releases and install the dependencies separately:

```bash
apt install ripgrep pandoc poppler-utils ffmpeg
```

Homebrew and MacPorts are also supported, and the README suggests adding the optional extractors after the main install:

```bash
brew install rga
brew install pandoc poppler ffmpeg
```

On Windows, chocolatey and scoop are described as the only supported download methods, because a manually downloaded binary does not bring the dependencies such as pdftotext from poppler. `choco install ripgrep-all` or `scoop install rga` are the two commands given.

If you prefer to build it, the README states that rga compiles with stable Rust v1.75.0 or later, and gives this sequence:

```bash
apt install build-essential pandoc poppler-utils ffmpeg ripgrep cargo
cargo install --locked ripgrep_all
rga --version
```

The crate is named `ripgrep_all` on crates.io even though the binary is `rga`. The README notes that rga looks for every binary it calls in `$PATH` and in its own directory, which is the fact to check first when an adapter silently does nothing.

For a first real search, point it at a directory and pass the pattern the same way you would to ripgrep. The usage line is `rga [RGA OPTIONS] [RG OPTIONS] PATTERN [PATH ...]`, so ripgrep's own flags still apply. To see what is available on your machine, run the adapter listing:

```bash
rga --rga-list-adapters
```

That output tells you which adapters loaded. If pandoc is missing, the docx and epub entries are effectively dead weight, and the search will simply skip those files rather than fail loudly.

## Where rga breaks down

The dependency chain is the main limitation. rga is a wrapper, and every adapter that matters calls an external program. A machine with rga installed but without poppler-utils will not find anything in a PDF. The README is explicit about the Windows case: downloading the release binary manually means you do not get the dependencies, which is why chocolatey and scoop are the only supported paths there.

Extension-based matching is the second weak point, and the project acknowledges it. Files named with the wrong extension, or with no extension at all, will not reach the right adapter unless you pay for `--rga-accurate`, which reads magic bytes and is described as slower. Even then, detection only inspects the first 8KiB, so a format whose signature sits later in the file can still be misclassified.

Caching is a third thing to think about. Extracted text is stored in a database under the user cache directory, and the README documents no invalidation policy or size ceiling beyond the phrase "if it is small enough". If you search sensitive documents, that cache is plain extracted text sitting on disk, and `--rga-no-cache` is the documented way to avoid it. Finally, this is the wrong tool for structured querying. The sqlite adapter renders a database as plain text so a regex can match it; if you actually want to run SQL, use sqlite3 directly.

## rga versus plain ripgrep and a manual extract pipeline

The obvious alternative is ripgrep plus a shell loop: find the archives, extract them to a temporary directory, then grep the result. That approach has no external extractor dependency beyond the tools you already chose, and you control exactly what gets extracted and where. What it loses is the recursion and the caching. rga descends into nested containers in one pass and reuses extracted text across searches, which a hand-rolled pipeline has to reimplement with its own cache directory and its own cleanup logic.

The second alternative is a document search system built on an index, such as a full-text engine that ingests PDFs and Office files ahead of time. That trades the on-demand model for a build step and a store to maintain, and it gives you ranked results rather than line matches. rga stays in the ripgrep mental model: regex in, matching lines out, no index to keep in sync. If your corpus is small and changes often, that is the better fit. If it is large and stable, the index approach usually wins on query latency.

## Maintenance, licensing and upgrade cost

The repository is not archived and the last push was on 2026-09-17, four days before this writing. Releases are infrequent but real: v0.10.10 on 2025-11-09, preceded by v0.10.9 and v0.10.8 on 2025-05-11. The Cargo manifest declares edition 2024 and a minimum of stable Rust v1.75.0, so building from source requires a reasonably current toolchain.

The licence is the part that needs a careful read. The README does not state a licence, and the repository metadata reports NOASSERTION, but Cargo.toml declares `license = "AGPL-3.0-or-later"`. AGPL is a strong copyleft licence, and its network clause is broader than GPL's. If you are considering bundling rga into a service or shipping it inside a product, that declaration is the thing to take to whoever handles licensing on your side; this is a description of what the file says, not legal advice. For internal command-line use the practical effect is usually nil, but the distinction matters before redistribution.

Upgrade cost is low in normal use. Because the adapters shell out to pandoc, poppler and ffmpeg, an upgrade to rga does not usually force those tools to change, and the cache format is the main internal detail that could shift between versions. The CHANGELOG.md at the repository root is where release notes live.

## Conclusion

Adopt rga if you already use ripgrep and regularly need to grep through PDFs, docx, epub, sqlite databases or nested archives, and you are willing to install pandoc, poppler-utils and ffmpeg alongside it. Do not adopt it if you need a single self-contained binary, if your target formats are not covered by an adapter, or if you cannot install external tools on the machine. Before rolling it out, run rga --rga-list-adapters on the target machine to confirm which adapters resolved, and check where the cache database lands, since --rga-no-cache is the only way to turn it off.

## FAQ

### How does ripgrep-all work?

rga wraps ripgrep and routes files through adapters that convert each format to plain text, then lets ripgrep match lines in the result. The zip, tar and decompress adapters recurse into containers, so a PDF inside a zip can still be searched in one pass.

### Is ripgrep-all better than grep?

They solve different problems. ripgrep is a fast line-oriented search tool over text files, while rga wraps ripgrep and adds adapters so the same search reaches PDFs, docx, epub, sqlite and archives that grep would treat as opaque bytes.

### Can I install ripgrep-all using Homebrew?

Yes. The README gives `brew install rga`, and separately suggests `brew install pandoc poppler ffmpeg` for the dependencies, which it describes as not strictly necessary but very useful.

### How can I use ripgrep-all to find files?

Pass a pattern and a path the same way you would to ripgrep, using the form `rga [RGA OPTIONS] [RG OPTIONS] PATTERN [PATH ...]`. ripgrep's own options still apply alongside rga's.

## Sources

- [Issues](https://github.com/phiresky/ripgrep-all/issues)
- [phiresky/ripgrep-all on GitHub](https://github.com/phiresky/ripgrep-all)
- [README](https://github.com/phiresky/ripgrep-all/blob/master/README.md)
- [Releases](https://github.com/phiresky/ripgrep-all/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/phiresky-ripgrep-all
