CLI tool
wireservice/csvkit avatar
wireservice/csvkit

csvkit: A Command-Line Suite for Converting, Filtering, and Analyzing CSV Files

A suite of utilities for converting to and working with CSV, the king of tabular file formats.

6,415 stars693 forksPythonMIT

At a glance

What is it?
csvkit is a set of Python command-line tools that converts spreadsheets and database exports into CSV and then manipulates them with Unix-pipe-friendly utilities, targeting data journalists, analysts, and developers who work with tabular data in the shell.
Who is it for?
Data analysts, journalists, and developers who work with tabular data in the terminal and need to convert spreadsheets, run SQL-like queries, or inspect CSV structure without loading Python interactively should install csvkit. Projects that need to process files too large to fit in memory, or that need streaming CSV handling, should look at Miller or xsv instead, since csvkit loads files into memory via the agate library.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 10 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What csvkit is and who it targets

csvkit grew out of a practical need in data journalism: reporters and analysts routinely receive data as Excel spreadsheets, Access databases, or fixed-width text files, and need a reliable way to normalize them into CSV for further analysis. The project, maintained by wireservice, describes itself as inspired by pdftk, GDAL, and the original csvcut tool by Joe Germuska and Aaron Bycoffe.

The library agate sits at the core of csvkit's processing model. Agate handles type inference, encoding, and tabular operations. This means csvkit understands that a column containing 2024-01-15 is a date, that a column with 1,234 is a number, and that a column with 'yes' and 'no' is boolean, without you having to declare schemas. The convenience is real, but it has a cost in memory.

The toolset: what each command does

csvkit ships a family of composable command-line tools. The conversion tools accept formats beyond CSV and emit normalized CSV: in2csv handles Excel (.xlsx, .xls), dBASE (.dbf), and other formats. sql2csv runs a SQL query against a database and writes the results as CSV.

The manipulation tools operate on CSV files: csvcut selects or reorders columns, csvgrep filters rows by regex or exact match, csvsort sorts rows by one or more columns, csvstack concatenates multiple CSV files vertically, and csvjoin performs SQL-like joins between two files. Analysis tools include csvstat, which computes descriptive statistics for every column, and csvlook, which renders a CSV as a Markdown table for quick inspection.

csvclean checks for structural errors and can fix common problems. csvformat converts between CSV dialects, including TSV. All tools read from stdin and write to stdout, making them composable in shell pipelines.

Installing csvkit and running a first conversion

csvkit version 2.2.0 requires Python 3.10 or newer and installs from PyPI. The Dockerfile in the repository shows the minimal install:

dockerfile
FROM python:alpine

COPY csvkit csvkit
COPY man man
COPY README.rst pyproject.toml .

RUN pip install --no-cache-dir .

For a local install, pip reads the project name directly from the package index. The project's pyproject.toml lists its name as csvkit with version 2.2.0. The optional zstandard compression support installs with the extras flag. The documentation lives at csvkit.readthedocs.io and covers every command with examples from the examples/ directory in the repository, which contains sample CSV, Excel, TSV, and compressed files for testing.

How csvkit handles Excel, DBF, and SQL sources

The agate-excel, agate-dbf, and agate-sql packages extend csvkit's in2csv command. agate-excel uses openpyxl for .xlsx files and xlrd for .xls files. agate-dbf handles dBASE format, common in GIS and legacy government data exports. agate-sql uses SQLAlchemy to connect to any database with a supported driver, letting you export a table or a query result directly to CSV.

This means csvkit's dependency tree includes SQLAlchemy for the SQL-to-CSV path. If you only use csvkit for file-based operations, SQLAlchemy is still installed but unused. The agate library handles encoding detection, which reduces problems with Windows-1252 encoded files common in older government data exports.

The examples/ directory in the repository includes iris.csv, join_a.csv, join_b.csv, and several files designed to test csvclean's error detection, such as bad.csv and blanks.csv.

csvkit vs Miller and xsv: three tools with different audiences

Miller (mlr) and xsv are the two most common alternatives mentioned alongside csvkit. Miller is a streaming, record-oriented data processing tool written in Go; it handles files larger than memory because it processes one record at a time rather than loading the whole file. Miller's query language is its own DSL, which has a steeper learning curve than csvkit's separate named commands.

xsv is a Rust-based CSV toolkit focused on speed and large file handling. It reads files in a single pass with minimal memory overhead. xsv does not include SQL integration or Excel conversion; it is narrowly focused on CSV manipulation at high throughput.

csvkit's advantage is breadth: conversion from non-CSV formats, SQL integration, and descriptive statistics in one installation. Its limitation is memory: large files that do not fit in RAM will cause failures or slow performance. For files under a few hundred megabytes on a modern machine, csvkit's convenience usually outweighs the memory cost.

Limitations: memory usage and no streaming support

agate loads the entire dataset into memory before any operation can run. A 2GB CSV file requires 2GB or more of free RAM to process. csvkit does not offer a streaming mode. This is a deliberate design trade-off: full in-memory access makes type inference, sorting, and joining straightforward to implement, but it caps the practical file size.

Type inference can also misfire. A column containing values like 01, 02, 10 might be inferred as integers, stripping the leading zeros. A column containing dates in an ambiguous format may be parsed incorrectly. csvclean can detect some of these problems, but it cannot fix type inference errors; that requires explicit options passed to the specific command.

The project requires Python 3.10 or newer since version 2.2.0. Environments locked to Python 3.8 or 3.9 cannot use the current release without upgrading. The README does not document a maintenance branch for older Python versions.

Maintenance and licensing

The repository was last pushed on 2026-09-21, with active changes to the CI configuration and documentation. The project version in pyproject.toml is 2.2.0 and the license is MIT.

The project is maintained by Christopher Groskopf and James McKinney, listed as authors in pyproject.toml. The development status classifier in pyproject.toml reads 'Production/Stable', indicating the maintainers consider it suitable for production use. Issue tracking is on GitHub at wireservice/csvkit. The documentation at csvkit.readthedocs.io is built with .readthedocs.yaml in the repository root.

The repository includes man pages in the man/ directory and a Dockerfile for containerized use. The MIT license allows use in commercial projects without restriction. The project also has a schemas repository (wireservice/ffs) linked from the README for users who need to work with field-level metadata alongside their CSVs.

Editorial conclusion

Data analysts, journalists, and developers who work with tabular data in the terminal and need to convert spreadsheets, run SQL-like queries, or inspect CSV structure without loading Python interactively should install csvkit. Projects that need to process files too large to fit in memory, or that need streaming CSV handling, should look at Miller or xsv instead, since csvkit loads files into memory via the agate library. Before installing, confirm Python 3.10 or newer is available; the project dropped support for older Python versions in the 2.x series.

Frequently asked questions

How do I install csvkit?

csvkit is available on PyPI. Install it with pip from the command line. It requires Python 3.10 or newer. The optional zstandard compression support can be added with the extras syntax.

How do I use csvkit?

csvkit provides separate commands for each task: in2csv converts Excel and other formats to CSV, csvcut selects columns, csvgrep filters rows, csvstat shows statistics, and csvjoin joins two files. All commands read from stdin and write to stdout, so they compose in shell pipelines. The documentation at csvkit.readthedocs.io has worked examples.

How does csvkit compare to Miller?

Miller processes records one at a time in a streaming fashion and can handle files larger than available RAM. csvkit loads the full dataset into memory via agate, making it unsuitable for very large files but giving it richer format conversion and SQL integration through in2csv and sql2csv.

How does xsv compare to csvkit?

xsv is a Rust-based CSV toolkit that prioritizes speed and handles large files with low memory usage. csvkit covers more ground: it converts from Excel, dBASE, and SQL sources, and includes statistical analysis via csvstat. xsv does not include those conversion tools.

What are alternatives to csvkit?

Miller (mlr) is a streaming alternative for large files with its own query language. xsv is a fast, low-memory Rust tool for CSV manipulation. For Python users who prefer an interactive environment, pandas covers most of the same operations inside a Jupyter notebook.

Official sources

  1. Issues
  2. License: MIT
  3. Project website
  4. README
  5. wireservice/csvkit on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/wireservice-csvkit.svg)](https://hysenlabs.com/projects/wireservice-csvkit)