Model or dataset
idank/explainshell avatar
idank/explainshell

idank/explainshell: matching command-line arguments to manpage help text

match command-line arguments to their help text

14,263 stars852 forksPythonGPL-3.0

At a glance

What is it?
explainshell parses a shell command into an AST and maps each argument to the help text that documents it. This is a review of the self-hosted Python project behind explainshell.com, its SQLite storage model, and the LLM-assisted extraction pipeline.
Who is it for?
Adopt explainshell if you want a self-hosted or offline way to resolve shell arguments to real manpage text, and if you are comfortable running its manager CLI to build or download a SQLite database. Skip it if you need coverage of commands whose manpages are not in the archive, or if you want a drop-in library rather than a Flask application.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 35 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem: a command line you cannot read

A long shell invocation is a dense sequence of flags, values and sub-commands. To understand it you open a manpage, search for the flag, and repeat that for every argument. The README frames the project as a tool that parses manpages, extracts options, and explains a given command line by matching each argument to the relevant help text in the manpage. The target user is someone reading an unfamiliar command, whether in a tutorial, a shell history, or a colleague's script. It is not a shell, not a linter, and not a replacement for `man`. It answers one question: what does this specific token mean in this specific program. The project is GPL-3.0 and the last push to the master branch was on 2026-08-26.

Four components and one SQLite file

The architecture is described as four pieces. A manpage reader (manpage.py) extracts metadata: name, synopsis and aliases. An options extractor under extraction/ parses roff macros or uses an LLM to extract options. A storage backend (store.py) saves processed manpages to SQLite. A matcher (matcher.py) walks the command's AST, parsed by bashlex, and contextually matches each node to help text.

At query time the flow is: parse the query into an AST, visit interesting nodes such as command nodes and shell nodes for `|` and `&&`, check whether the current program is known, then walk the remaining tokens against the list of known options. Matches are rendered with Flask.

The storage model is a single `explainshell.db` with three tables. `manpages` holds zlib-compressed source text, typically markdown produced by `mandoc -T markdown`. `parsed_manpages` holds extracted options as a JSON list, plus synopsis, aliases and behavioral flags. `mappings` maps command names to `parsed_manpages` rows many-to-one, with a score for preference, so one manpage can have several mappings, one per alias and one per sub-command form such as `git commit`. The `source` path, formatted `distro/release/section/name.section.gz`, is the primary key across the first two tables and doubles as a namespace, so queries can be scoped to a distro and release by filtering on the path prefix.

Running explainshell locally and explaining your first command

The README gives a local setup. It clones the repository, creates a virtualenv, installs the dev requirements, downloads the live database, and starts the server on port 5000.

bash
git clone https://github.com/idank/explainshell.git
cd explainshell
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements-dev.txt
make download-latest-db
make serve

After `make serve` the README says to open http://localhost:5000. The live database is saved in a GitHub release, which is why the download step exists: you do not have to parse manpages yourself before trying the web interface. The runtime dependencies in requirements.txt are Flask, gunicorn, bashlex, pydantic, cmarkgfm, humanize and cachetools, so the web application itself does not need an LLM provider key. According to .env.example, the provider variables are used by offline extraction tooling, not by the web app.

If you would rather build a database from a manpage, the manager CLI takes a database path from the `DB_PATH` environment variable or a `--db` flag. Commands that do not need a database, such as `extract --dry-run` and `diff extractors`, work without one.

bash
python -m explainshell.manager --db test.db extract --mode llm:codex/gpt-5.6-sol/medium manpages/ubuntu/26.04/1/tar.1.gz

The README states this calls out to `codex exec` to extract the manpage and writes the result to test.db. The `--mode` flag selects the extraction strategy; an API key path is also documented, for example `--mode llm:openai/gpt-5-mini`, with keys configured per .env.example. To inspect what landed in the database, the manager has `show stats`, `show manpage <name>`, `show distros` and `db-check`.

Why the LLM extractor returns line ranges

The extraction design is the most interesting decision in the repository, and it is also the one most likely to be misread. In `llm:<provider/model>` mode the manpage text is converted to markdown with `mandoc -T markdown` and sent to a model. The README says the LLM returns line ranges into the source text, not generated descriptions, so hallucinations are structurally impossible: the actual help text is always sliced from the original manpage. That is a real constraint on the model's output, not a prompt request, and it means a bad extraction produces a wrong range rather than invented prose.

The CLI reflects that this is a batch pipeline rather than a one-shot tool. `--overwrite` re-processes existing entries, and with it `--filter-db <spec>` re-extracts only rows whose stored extractor matches a spec, repeatable to match several. `--dry-run` extracts without writing. `-j <N>` sets parallel workers. `--batch <N>` uses a provider batch API for LLM modes including `gemini/`, `openai/` and `azure/`. `--small-only` and `--large-only` partition the corpus at roughly 2 KB gz so a cheap model handles small pages and a capable one handles the rest. The README's two-pass example runs a cheap model with `--small-only` first, then a capable model with `--large-only`, noting that already-stored small pages are skipped. `diff db` compares a fresh extraction against the database, and `diff extractors` compares two extractors head-to-head with the `A..B` syntax.

Where explainshell stops being the right tool

Coverage is bounded by the archive. The README states manpages come from known archives, currently the Ubuntu archive and manned.org, and that all sources are committed to explainshell-manpages, a git submodule. `make ubuntu-archive UBUNTU_RELEASE=resolute` and `make arch-archive` generate gzipped manpages under `manpages/<distro>/<release>/`, with the note that Arch only has a latest archive. A command whose manpage is missing from those sources has no help text to match, and no amount of matcher logic changes that. This is the failure mode to plan for: a query that returns partial or no explanation because the program is unknown, not because the parsing failed.

The second boundary is that this is an application, not a library. The README describes a web interface and a Flask rendering step; there is no documented Python API for embedding the matcher in another tool. If you want argument explanation inside an editor or a CI job, you are integrating against a running service or reusing the components on your own terms, and the README does not document that path. The third is operational: the database is a build artifact. It comes from a GitHub release or from your own extraction run, and refreshing it means re-running extraction, which for LLM modes means provider credentials, rate limits and cost. The README does not document rollback to a previous database version, so keeping the previous file before replacing it is on you.

Alternatives and the difference in approach

The closest alternative in the search results people use is the system `man` command, and the difference is one of addressing. `man tar` gives you the whole page and leaves you to find the flag; explainshell parses the command into an AST and returns per-token matches, so it must know the program and the option list before it can answer. That is a strictly larger prerequisite: `man` works on any installed manpage, while explainshell needs the page in its database and a mapping for the command name.

The second comparison is shell completion and `--help` output. Those are generated by the tool itself and reflect the version you have installed; explainshell reads manpage text from an archive, so the explanation is only as current as the archive and the database build. If a flag was added after the manpage snapshot, explainshell will not know it. The trade is that explainshell can explain a command you do not have installed, which neither completion nor `--help` can do, and it can explain shell syntax such as pipes and `&&` because it parses the AST rather than the program's own option table.

Maintenance, upgrades and the GPL-3.0 boundary

The repository is not archived, and the last push was on 2026-08-26. The most recent release listed is db-latest from 2026-03-05, which is the database artifact rather than a code version, so the practical upgrade path for the web service is a code pull plus a fresh `make download-latest-db`. The Makefile shows the production image build resolving the database asset name through `gh api` against the `db-latest` release and failing with a message if it cannot resolve it, which means the release tag and asset naming are load-bearing for deployment, not incidental.

Upgrade cost concentrates in extraction, not in the Flask app. Changing extractor mode, model or provider means re-running extraction over the corpus, and `--filter-db` exists precisely so you can re-extract only the rows produced by one extractor. The eval harness under tests/evals/llm runs the extractor on a corpus file and writes a summary plus per-page artifacts, with runs saved under timestamped directories, so extractor changes can be compared before they reach the live database. The test entry point is `make tests-all`, described as lint plus unit tests plus e2e.

On licensing: the project is GPL-3.0. If you modify it and distribute it, or run a modified version as a network service, the licence terms apply to that distribution or deployment. The manpage text itself comes from distribution archives and manned.org, and the README does not state the licence of the archived manpage sources, so that is a question to settle before republishing extracted text. This is a description of what the repository states, not legal advice.

Editorial conclusion

Adopt explainshell if you want a self-hosted or offline way to resolve shell arguments to real manpage text, and if you are comfortable running its manager CLI to build or download a SQLite database. Skip it if you need coverage of commands whose manpages are not in the archive, or if you want a drop-in library rather than a Flask application. Before committing, run `python -m explainshell.manager show manpage tar` against the downloaded database to confirm the commands you care about are present and mapped.

Frequently asked questions

Is explainshell down?

The repository does not document the availability of the hosted service, so there is no status information to check here. The README does give a local path: clone the repository, run `make download-latest-db`, then `make serve` and open http://localhost:5000.

Can you explain the Linux shell and what it does?

That is a general shell question rather than something this repository documents. What explainshell does with shell input is parse it into an AST with bashlex, then visit command nodes and shell-related nodes such as `|` and `&&`, and match the remaining tokens against known options.

Is shell scripting hard to learn?

The README does not discuss learning shell scripting. It describes explainshell as a tool that parses manpages, extracts options, and matches each argument of a given command line to the relevant help text, which is aimed at reading commands rather than writing scripts.

How do I check my terminal shell?

explainshell does not report which shell you are running. It parses the command line you give it, including shell operators, using bashlex, and explains the arguments; the README does not describe any shell-detection feature.

What is shell used for?

The README does not cover that topic. Its scope is narrower: explaining a specific command line by matching each argument to help text extracted from a manpage, with shell nodes such as pipes and `&&` handled as part of the parsed AST.

Official sources

  1. idank/explainshell on GitHub
  2. Issues
  3. License: GPL-3.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/idank-explainshell.svg)](https://hysenlabs.com/projects/idank-explainshell)