# Miller (mlr): name-indexed data processing for CSV, TSV and JSON

> Miller is a single Go binary that treats records as insertion-ordered hash maps instead of positional fields. It is for engineers who reshape CSV, TSV, JSON and log data on the command line, and it stays streaming for most operations.

**johnkerl/miller** — Miller is like awk, sed, cut, join, and sort for name-indexed data such as CSV, TSV, and tabular JSON

- Repository: https://github.com/johnkerl/miller
- Website: https://miller.readthedocs.io
- Stars: 10,031 · Forks: 244
- Language: Go
- License: NOASSERTION
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/johnkerl-miller

## The problem Miller solves: records with names, not column numbers

The classic Unix toolkit indexes fields by integer. cut -f3, awk '{print $7}' and sort -k2 all assume you know where a column sits and that it stays there. That assumption breaks the moment a producer adds a column, reorders a header, or emits records with different schemas in the same file. Miller's stated premise is that its natural data structure is the insertion-ordered hash map, while the Unix tools' natural structure is the array. If you have ever rewritten a cut pipeline because someone inserted a field, that sentence is the whole pitch.

The audience is specific. People cleaning exports before loading them into pandas or R, people pulling fields out of log files and emitting CSV, people doing one-off queries against disk files without standing up a database. The README lists data cleaning, data reduction, statistical reporting, devops, system administration, log-file processing, format conversion and database-query post-processing as intended uses. It is not a dataframe library and it does not pretend to be one. It complements R and pandas by preparing data rather than analysing it.

The second premise matters as much as the first: record heterogeneity. Records with different field names can be interleaved in one stream, which is normal in JSON Lines and rare in CSV. A tool built on positional fields cannot express that without padding. A tool built on hash maps can.

## How mlr works: streaming records through a verb chain

A Miller invocation is a chain of verbs separated by then. Each verb consumes records and emits records. Input is parsed into records by a format reader, verbs transform them, and an output format writer serialises them. Because the record is a hash map, a verb like cut -f amount refers to a name, and the name is resolved per record rather than per file position.

The README states that Miller is streaming: most operations need only a single record in memory at a time rather than ingesting all input before producing output. That is the property that lets it run in tail -f contexts and over files larger than RAM. Deep-retention verbs are named explicitly: sort, tac and stats1 retain only as much data as needed. That phrasing is doing real work. sort by definition cannot emit its first record until it has seen its last, and stats1 must hold accumulators. So the streaming claim is a claim about most verbs, not all of them, and the exceptions are the ones you would guess.

Processing is format-aware. The README gives CSV sort and tac as examples that keep header lines first. A positional tool would sort the header row into the data. There is also a format conversion role: the same record stream can be read as CSV and written as JSON, which is how Miller gets used as a pre-processor for tools that only accept one shape.

The implementation is modern Go with zero runtime dependencies, distributed as a single binary. The go.mod module path is github.com/johnkerl/miller/v6 and the executable is mlr, a name the repository comments say predates the Go port. The build target in the Makefile points at github.com/johnkerl/miller/v6/cmd/mlr, and the comment explains why: with ./mlr.go at the root the binary would be named miller, so the entry point lives under cmd/mlr instead.

## Installing Miller and running a first conversion

The README's installation table gives package-manager commands per platform. On Debian or Ubuntu the package is miller, installed with apt-get. On Fedora or RHEL-alikes, yum. On macOS, brew install miller or port install miller. On Windows, choco install miller, winget install Miller.Miller, or scoop install main/miller. Snap is also listed. The README notes that long-term-support releases will likely be on older versions, which is the practical caveat: a distribution package may lag the upstream release.

```bash
apt-get install miller
```

After installation the executable is mlr. The README does not show a version-check command, but the binary name is what every example uses, so that is the thing to confirm is on your PATH before going further.

Building from source is also documented, and the Makefile makes it a two-step. The default target builds the binary in the repository root.

```bash
make build
```

The Makefile prints that the executable is ./mlr (or .\mlr.exe on Windows), and points at make check for tests. Installing to a non-default prefix is handled by configure before make install; the default prefix is /usr/local and the install directory is /usr/local/bin.

```bash
./configure --prefix=/your/install/path
make install
```

For a first real use, the shape to internalise is verb chains joined by then. The README describes adding fields that are functions of existing fields, dropping fields, sorting, aggregating statistically and pretty-printing. The repository ships a data/ directory and a test/ directory, and the documentation links a "Miller in 10 minutes" page, which is where the worked examples live. The README itself does not spell out a full command with sample input and output, so treat the linked tutorial as the source for syntax rather than guessing at verb flags. A practical first exercise is to take one of your own CSV exports and run it through a chain that selects a few named fields and writes JSON, then compare the output against the same selection done with cut and paste.

## Where Miller is the wrong tool

The streaming guarantee has a boundary, and it is the boundary most people hit. The README names sort, tac and stats1 as the operations that require deeper retention. If your pipeline is a global sort of a hundred-gigabyte file, Miller will hold what it needs to hold, and the claim that you can operate on files larger than RAM applies "whenever functionally possible". That qualifier is honest and worth reading twice. A streaming tool cannot make a global sort free.

Statistical work is deliberately shallow. The README says you can do basic statistics entirely in Miller, and that it complements R, pandas and similar tools. It does not claim regression, modelling or plotting. The go.mod does pull in github.com/kshedden/statmodel, but the README's own framing keeps the statistical surface at "basic", and the release notes for v6.21.0 mention new verbs and DSL features rather than analytical depth.

Joins are the other soft spot. The description compares Miller to join, and Miller does have join verbs, but there is no persistent index and no query planner. If you are joining several large tables repeatedly, a database is the right answer, and the README says as much by positioning Miller as complementary to SQL databases rather than a replacement. The one-off query use case it does claim is narrow: querying data in disk files in a hurry without setup.

Finally, schema enforcement. Record heterogeneity is a feature here, which means Miller will not stop you from processing a file where half the records are missing a field. If you need validation and rejection, that is a separate step.

## Miller against awk, and against pandas

The nearest comparison is awk, and the difference is not speed, it is the unit of address. awk addresses fields positionally with $1, $2, $3. Miller addresses them by name. For a fixed-format log line, awk is fine and often shorter. For a CSV with a header, the awk version has to either hard-code column numbers or parse the header itself, and it has to handle quoting. Miller's format-aware reader does that parsing for you, and the same verb chain works whether the input is CSV, TSV or JSON Lines. That format independence is the real distinction: an awk program is written against one shape, a Miller chain is written against field names.

The other comparison is pandas or R. Those load data into memory, give you a typed columnar frame, and give you a vast analytical surface. Miller does not build a frame and does not give you that surface. What it gives you instead is the ability to run over input that never fits in memory, with no runtime dependencies and no environment to set up. The README frames the relationship as complementary: use Miller to clean and prepare, then hand the result to the in-memory tool. If your dataset fits comfortably in RAM and you are doing real analysis, going straight to pandas is less work than learning a verb chain. Miller earns its place at the boundary, not in the middle.

## Maintenance, versioning and what the licence file leaves open

The repository is not archived. The last push was on 2026-09-21, and the most recent release listed is v6.21.0 on 2026-08-10, described as adding new verbs and DSL features. Before that, v6.20.2 on 2026-07-04 covered Miller and AI, bash/zsh tab-completion and a bytes datatype, and v6.19.0 on 2026-06-19 covered performance improvements. Three releases in roughly three months, with feature, tooling and performance work in different ones, is a reasonable signal about the pace of change.

Upgrade cost is mostly a matter of where you get it. Distribution packages are the easy path and the README warns that long-term-support releases will likely be on older versions, so a verb added in a recent release may simply not exist in your apt or yum package. Building from source removes that lag and requires a Go toolchain; go.mod declares go 1.26.0, and the module has both direct and indirect dependencies, so a source build is not dependency-free even though the resulting binary is. Cross-version script compatibility is the other cost to budget for: a verb chain is a script, and verbs and DSL features do change between minor versions, so pinning a version for anything you run unattended is the low-effort hedge.

On licensing, the repository has a LICENSE.txt at the top level, but the metadata reported for this repository is NOASSERTION, meaning the licence could not be automatically identified. The README does not state the licence. That is a gap you have to close yourself by reading LICENSE.txt before redistributing the binary or bundling it into a product. Nothing here is legal advice; the point is that the machine-readable signal is absent and the file is the authority.

## Conclusion

Adopt Miller if your daily work is reshaping delimited or JSON records in shell pipelines and you want named fields, format conversion and streaming in one binary. Skip it if you need SQL joins across many tables, a persistent index, or dataframe-style columnar analytics; pandas, R and a real database are the better fit there. Before committing, verify two things yourself: that your distribution's package version is recent enough for the verbs you plan to use, and that the operations you rely on most are the streaming ones, since sort, tac and stats1 retain data. The licence file is LICENSE.txt, and the README does not state which licence it contains, so read it before you redistribute the binary.

## FAQ

### How do I install Miller (mlr)?

The README gives package-manager commands per platform: apt-get install miller or yum install miller on Linux, brew install miller or port install miller on macOS, and choco install miller, winget install Miller.Miller or scoop install main/miller on Windows. Snap is also listed. The README notes that long-term-support releases will likely be on older versions.

### What is the Miller command-line tool used for?

The README describes it as like awk, sed, cut, join and sort for name-indexed data such as CSV, TSV, JSON and JSON Lines, and lists data cleaning, data reduction, statistical reporting, devops, log-file processing, format conversion and database-query post-processing. It works on named fields rather than positional indices.

### Why is the Miller executable called mlr and not miller?

The comment in go.mod explains it: the executable has been named mlr for many years, predating the Go port, and if the entry point were ./mlr.go at the repository root then a plain go build would produce a binary named miller. The entry point therefore lives at cmd/mlr.

### Does Miller need to load the whole file into memory?

The README states that most operations need only a single record in memory at a time, so you can operate on files larger than your available RAM whenever functionally possible. It names sort, tac and stats1 as the operations that require deeper retention, and says those retain only as much data as needed.

### What licence is Miller released under?

The repository has a LICENSE.txt at the top level, but the reported licence metadata is NOASSERTION and the README does not state the licence. Read LICENSE.txt directly before redistributing the binary.

## Sources

- [Issues](https://github.com/johnkerl/miller/issues)
- [johnkerl/miller on GitHub](https://github.com/johnkerl/miller)
- [Project website](https://miller.readthedocs.io)
- [README](https://github.com/johnkerl/miller/blob/main/README.md)
- [Releases](https://github.com/johnkerl/miller/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/johnkerl-miller
