whatisit-nl2sh: a local natural-language to shell command generator
Local natural-language-to-shell command generator. A 941 MB fine-tuned Qwen2.5-Coder-1.5B running on CPU in ~1s.
At a glance
- What is it?
- whatisit turns a plain-English request into a shell command using a 941 MB fine-tuned Qwen2.5-Coder-1.5B that runs on CPU through llama.cpp, with no network by default. It is a good fit for offline machines and restricted environments, and the wrong tool when you need an explanation of a command or a guarantee that a generated line is safe to run.
- Who is it for?
- whatisit suits engineers who work on offline or restricted machines and want command suggestions without sending shell context to a cloud API. It is a poor fit if you need the model to explain a command or if you expect a safety guarantee: the README states the checker cannot see everything, so review before you confirm.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap whatisit-nl2sh fills: shell syntax you half remember
Most people know what they want the shell to do and not the flags that do it. The usual fix is a search engine or a cloud chat endpoint, and both send your request off the machine. whatisit-nl2sh takes the other route. You type the request as plain arguments, and a local model returns a command. The README's own examples are ordinary sysadmin chores: finding files bigger than 100MB, taring a logs directory, deleting every .pyc file under a tree.
The target user is someone on a machine where piping shell context to a cloud API is not allowed, or someone who simply works offline. The README says the model is 941 MB and runs through llama.cpp, and that by default nothing you type leaves the machine. That constraint is the whole reason the project exists. A cloud model would answer these questions better and cheaper to host; it would also require a network round trip and a data-sharing decision that some environments forbid.
There is a second, quieter use case: recall. You know the command exists and cannot reconstruct the flag order. Greedy decoding at temperature 0 means the same question returns the same command every time, which is what you want from a lookup tool and not what you want from a creative one.
How the local model, the CLI and llama.cpp fit together
Three pieces have to be present: the CLI, the GGUF model file, and a llama.cpp build that runs it. The CLI has no dependencies of its own, and the README states the Python requirement is 3.9+ on Linux or macOS.
On the first call the CLI starts a small llama.cpp server and leaves it resident, so later calls skip model loading. That is the difference between the documented cold start of 2.1s and a warm median latency of 0.59s over 12 queries. The README reports those numbers from an Intel i5-11320H with 4 cores using 4 threads, along with 39.5 tok/s generation and 1.6 GB resident memory. Treat them as one machine's figures, not a specification.
The model itself is a LoRA fine-tune of Qwen2.5-Coder-1.5B-Instruct, merged into the base weights in bf16, converted to f16 GGUF, then quantized to Q4_K_M. The README documents the training run: LoRA rank 32, alpha 64, dropout 0.05, all linear target modules, 2e-4 learning rate with cosine schedule and 3% warmup, two epochs, batch 16 with 2x gradient accumulation, sequence length 512, bf16 at seed 42, on one A100 80GB for about an hour over 125,770 NL/command pairs. Because the model is a 1.5B at Q4_K_M, the README's own framing is that thread count matters more than the CPU model.
There is a fallback path: the README notes you need llama-server, and llama-cli for the fallback. So the resident-server design has a slower route behind it when the server cannot be used.
Installing whatisit and running a first command
The quick path is two commands. pipx installs the CLI in an isolated environment, and setup works out what the machine needs, tells you the size of each download, asks before starting it, and then checks the files it fetched.
pipx install whatisit
whatisit setupOn Linux with a glibc older than 2.34, setup picks a compatibility build, because the README states the upstream llama.cpp binaries will not start there. pip install whatisit also works inside a virtualenv or on conda.
If you would rather assemble the pieces yourself, the README breaks it into the CLI, the 941 MB model, and a llama.cpp build. The model download uses the Hugging Face CLI:
pip install -U huggingface_hub
hf download ThorOdinson246/nl2sh-1.5b-Q4_K_M nl2sh-1.5b-Q4_K_M.gguf --local-dir .If hf is not found, your huggingface_hub predates the rename; upgrade it or use huggingface-cli download with the same arguments. Then point the CLI at both files and check the result:
whatisit setup --model ./nl2sh-1.5b-Q4_K_M.gguf --bin-dir /path/to/llama.cpp/bin
whatisit doctor--bin-dir is the directory containing llama-server, not the binary itself. doctor reports which of the three pieces is missing. A first real query needs no quoting:
whatisit list files changed in the last weekThe CLI prints the command. Nothing runs unless you pass -e and confirm at the prompt. For inline use, -q prints only the bare command, which the README shows inside command substitution:
cd "$(whatisit -q the directory holding the largest log file)"Nix users have a shorter route. The flake wires in llama.cpp from nixpkgs, so setup only fetches the model:
nix run github:ThorOdinson246/whatisit-nl2sh -- setupThe safety checker flags DANGER, and the README says what it cannot see
The most useful design decision here is that the tool does not run anything by default. You get a command, and execution requires both -e and a confirmation at the prompt. Anything flagged DANGER is never auto-run at all. The README's own example is a request to delete everything in the root directory, which produces rm -rf / with a DANGER marker reading "recursive force-delete of a critical path".
That is a pattern match against dangerous shapes, and the README is direct that the checker has limits: it points you to the Safety section for what gets flagged and what the checker cannot see. So the honest reading is that DANGER is a warning label, not a proof of safety. A command that is destructive in a way the checker does not model will come back unflagged. If you are the kind of user who would paste a suggestion straight into a root shell, this tool is not for you, and the README does not pretend otherwise.
The -e flow is the part worth adopting deliberately. Review, then run, is the intended loop, and -q exists for the cases where you already know the shape of the answer and want it inline. Those two modes have different risk profiles, and the README treats them separately.
Memory, idle timeouts and the resident server trade-off
Keeping the model resident is what makes the tool feel fast, and it is also its main cost. The README reports 1.6 GB of resident memory. On a laptop that is fine; on a small VM or inside a container with a tight memory limit it is the number that decides whether this is usable.
The README provides an escape hatch. A watchdog can stop the server after a period of inactivity and free the memory it was holding, and the next query pays a cold start instead. You can set it for one invocation or persist it in config:
whatisit --idle-timeout 300 show disk usage # this invocation only
whatisit config --set idle_timeout=300 # persist it (0 = never, default)The default is 0, meaning never. The deadline is re-armed on every query, including one still running, so a long generation will not be cut off mid-answer. That detail matters: a naive idle timer would kill the server during the slow first call, and the README states this one does not.
Thread count is the other knob, exposed through config. The README's framing is that for a 1.5B at Q4_K_M, how many threads you give it matters more than what they are in, which is a claim about the model size rather than about any particular CPU.
Remote endpoints: when local is the wrong default
whatisit can talk to any OpenAI-compatible endpoint instead of the local model, which covers a hosted API, Ollama, or a llama.cpp server you already run. Configuration is two settings plus an optional key:
whatisit config --set openai_base_url=http://127.0.0.1:8080/v1 openai_model=your-model-name
export WHATISIT_OPENAI_API_KEY=sk-... # optional; prefer env over the config fileopenai_model is required. Clearing the base URL, either through the environment variable WHATISIT_OPENAI_BASE_URL= or through openai_base_url=, returns to local mode. There are two more knobs: openai_timeout defaults to 120 seconds, and openai_max_tokens defaults to 512, which the README describes as being for reasoning models.
The README is explicit that this sends your request off the machine, that it is opt-in, and that whatisit prints a warning to stderr on every remote call. It also advises preferring https://. The recommendation to put the API key in the environment rather than the config file is the right instinct, and worth following.
This is also the clearest statement of where the local model is the wrong choice. A 1.5B quantized model will not match a hosted frontier model on unusual requests. If your environment permits remote calls and your requests are complex, switching to a remote endpoint is a one-line config change and probably the better answer. The local path is for when you cannot make that change.
What to check before adopting it, and the licence position
The project is Apache-2.0, and the repository carries a NOTICE file alongside the LICENSE. Apache-2.0 is permissive: it allows commercial and private use, modification and redistribution, with the usual conditions around preserving notices and stating changes. The NOTICE file is the thing to read rather than the licence identifier alone, since Apache-2.0 expects attribution notices to travel with redistributions. None of that is legal advice, and if you plan to redistribute a modified build, that is a question for your own counsel.
The model is a separate artifact from the code. The README documents that it is a LoRA fine-tune of Qwen2.5-Coder-1.5B-Instruct, which means the base model's own licence applies to the weights in addition to the repository's Apache-2.0 licence on the CLI. The README does not spell out the terms for the merged and quantized GGUF, so check the Hugging Face model page before you ship anything built on it.
On maintenance, the repository is not archived, and the last push was on 2026-09-15, two days before this writing. Releases are frequent: v0.4.0 on 2026-09-01, v0.3.0 on 2026-08-24, and v0.2.2 on 2026-08-14. That cadence is the useful signal, not any counter on the repository page.
Upgrade cost is low by design. The CLI has no dependencies of its own, so pipx upgrade whatisit moves the code without touching the model or the llama.cpp build. The pieces that can drift are the model file and the llama.cpp binaries, and setup re-checks the files it fetched. If you installed by hand, whatisit doctor is the command that tells you which piece is missing after an upgrade.
Editorial conclusion
whatisit suits engineers who work on offline or restricted machines and want command suggestions without sending shell context to a cloud API. It is a poor fit if you need the model to explain a command or if you expect a safety guarantee: the README states the checker cannot see everything, so review before you confirm. Verify first that whatisit setup picks a working llama.cpp build on your glibc version, and run whatisit doctor to see which of the three pieces is missing.
Frequently asked questions
What is whatisit-nl2sh?
It is a local natural-language-to-shell command generator: you describe what you want in plain English and it returns a shell command. The README states it uses a 941 MB fine-tuned Qwen2.5-Coder-1.5B running through llama.cpp, on CPU, with nothing leaving the machine by default.
How do I install whatisit-nl2sh?
Install the CLI with pipx install whatisit, then run whatisit setup. The setup command works out what your machine needs, reports the size of each download, asks before starting it, and checks the files it fetched. pip install whatisit also works inside a virtualenv or on conda.
Does whatisit-nl2sh run commands automatically?
No. The README states that nothing runs unless you pass -e and confirm at the prompt, and that anything flagged DANGER is never auto-run at all. The README also notes the checker cannot see everything, so the DANGER marker is a warning rather than a guarantee.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/thorodinson246-whatisit-nl2sh)