kardolus/chatgpt-cli: A Multi-Provider LLM Client With Agent Mode and MCP
ChatGPT CLI is a powerful, multi-provider command-line interface for working with modern LLMs. It supports OpenAI, Azure, Perplexity, LLaMA, and more, with features like streaming, interactive chat, prompt files, image/audio I/O, MCP tool calls, and an experimental agent mode for safe, multi-step automation.
At a glance
- What is it?
- The Go binary wraps OpenAI, Azure, Perplexity and LLaMA endpoints behind one config file, then adds an experimental ReAct agent with budgets and a workdir sandbox. The install is simple; the safety story is what deserves scrutiny.
- Who is it for?
- Adopt it if you want one CLI, one YAML config and thread-scoped history across several providers, and if you are willing to read the agent policy before enabling agent mode. Skip it if you need a stable automation surface rather than an experimental one, or if you cannot supply an API key.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap kardolus/chatgpt-cli fills between curl and a chat UI
Most people reach for an LLM either in a browser tab or through a handful of curl calls pasted into a shell script. The browser keeps history but cannot be piped into anything. The curl calls are stateless, so every request has to carry the whole conversation by hand. kardolus/chatgpt-cli sits between those two positions: a single Go binary that keeps thread-scoped history on disk, streams tokens to the terminal, and reads configuration from a layered YAML file.
The intended user is an engineer who already lives in a shell and wants the model available there. The README lists query mode for one-shot input and output, interactive mode for a conversational session, and streaming mode for real-time output, all against the same binary. The multi-provider angle matters because it means the same command can hit OpenAI, Azure, Perplexity, LLaMA, 302 AI or Atlas Cloud depending on which target is selected. That is a configuration problem rather than a code problem, and the project treats it that way.
Threads, a sliding context window, and where the tokens go
The mechanism behind the conversation memory is a thread identifier. Each unique thread carries its own history, so two terminals pointed at different threads do not see each other's messages. This is the same mental model as the OpenAI web interface, and it is the reason the CLI can be used for several unrelated tasks at once without context bleeding between them.
History cannot grow without limit, because every request is billed by the token. The README describes a sliding window: the chat history trims automatically while trying to preserve the context that still matters, and the size of that window is controlled by the context-window setting. That is a real trade-off rather than a detail. Set the window small and long conversations will lose earlier instructions; set it large and each turn costs more. The project also pulls in pkoukk/tiktoken-go, which is how it can reason about token counts locally rather than guessing.
Context does not have to come from the conversation alone. The README states that custom context can be piped in from any source, including local files, standard input, or another program. That turns the CLI into a component in a pipeline rather than only an interactive tool, and it is the feature most likely to be useful in a script.
Installing chatgpt-cli on macOS, Linux and Windows
There are two documented installation paths. Homebrew covers macOS, and direct downloads cover Apple Silicon, macOS Intel, Linux amd64, Linux arm64, Linux 386, FreeBSD amd64, FreeBSD arm64 and Windows amd64. The README does not describe a go install path, so building from source is a development concern rather than a documented install route.
On macOS the Homebrew route is the shortest. The README's installation section names the formula, and after it is installed the binary is invoked as chatgpt. The Getting Started section is the place to confirm which environment variable your chosen provider expects, and the configuration section covers the LLM-specific settings.
A first call is a plain prompt, and the answer should stream to the terminal rather than arriving all at once:
chatgpt "What is this photo?"To keep a session going instead of firing single queries, start interactive mode and the thread history will persist across turns. Model discovery uses the list flag:
chatgpt -lThat prints the models your configured provider reports. If the list is empty or the request fails, the problem is almost always the provider block in the config rather than the binary. The README also shows the binary being used inside a pipeline, which is the pattern to copy if you want to script it:
git status -vv | chatgpt -n -p ../prompts/create_git_diff_commit.mdAgent mode is the interesting part, and the part to be careful with
Agent mode runs multi-step tasks that can think, act and observe using tools such as shell commands, file operations and LLM reasoning. The README describes two shapes: an iterative ReAct loop and a Plan/Execute workflow. This is the feature that separates the project from a thin API wrapper, and it is also the feature labelled experimental.
The safety model has three parts, and they are worth reading before the first run. Budget limits cap time, steps and tokens. Policy enforcement decides which tools are allowed and which commands are denied. Workdir sandboxing constrains where file operations can reach. The README also documents a default policy, which means the tool ships with an opinion about what the agent may do rather than starting wide open.
My read is that the budget limits are the more valuable half. A ReAct loop that can call a shell is only as safe as its worst command, and no allowlist survives contact with a determined model. A step and token ceiling does something different: it bounds the blast radius in time and money even when the policy is imperfect. If you enable agent mode, set those ceilings first and treat the workdir sandbox as the second line of defence rather than the first.
MCP tool calls, sessions and how results re-enter the prompt
The CLI can call external MCP tools over HTTP(S) or STDIO, inject their results into the conversation context, and continue the prompt. The README's MCP section covers headers and authentication, session management, and how results are used.
Session handling is the detail that distinguishes a working MCP client from a demo. Stateful MCP servers expect a session identifier to be attached to follow-up calls, and the CLI initializes sessions, attaches the identifier, and renews it when the server reports it invalid. Without that renewal step, long tool-calling conversations fail partway through with errors that look like server faults. The repository also ships local test servers: the Makefile has mcp-http for a FastMCP HTTP server returning JSON and mcp-sse for a text/event-stream variant, both under test/mcp/http. Those are useful if you want to exercise the client without pointing it at a production server.
One thing to keep in mind is context cost. Every tool result is injected into the conversation, so a chatty MCP server will consume the same sliding window that your own messages share. There is no documented per-tool result cap in the README.
Images, audio and speech are provider-dependent
Beyond text, the CLI handles several media paths. The --image flag accepts a file or a URL, and images can be piped in directly, with the README giving the example of piping a macOS paste into a prompt. The --draw flag generates images and, combined with --image and --output, edits an existing one. Supported edit formats are PNG, JPEG and WebP, and generation requires an image-capable model such as gpt-image-1.
Audio splits into two features with different format support. Uploading audio with --audio works only with audio-capable models such as gpt-4o-audio-preview, and only .mp3 and .wav are accepted. Transcription via --transcribe uses the transcription endpoint and accepts a wider set: .mp3, .mp4, .mpeg, .mpga, .m4a, .wav and .webm. Text to speech uses --speak with --output, and the README shows chaining playback on macOS through afplay:
chatgpt --speak "convert this to audio" --output test.mp3 && afplay test.mp3The limitation is stated plainly in the README: image support may not be available for all models. The same logic applies to audio. These flags are not general capabilities of the client, they are capabilities of whichever model you have configured, and a mismatch produces a provider error rather than a helpful local message.
Where this CLI is the wrong tool, and what to use instead
The clearest limitation is the experimental label on agent mode. Multi-step automation that can run shell commands is exactly the kind of feature you do not want to depend on for unattended production work until it has settled. If your requirement is a deterministic batch job, a small script calling the provider's HTTP API directly gives you a surface you control completely, with no policy layer, no budget layer and no version churn between you and the request.
A second boundary is authentication. There is no documented path to use the tool without an API key, so anyone looking for a free client that reuses a consumer subscription will not find it here. The README's configuration sections all assume a provider credential.
The obvious alternative is OpenAI's own Codex CLI, which people search for alongside this project. The difference in approach is scope. Codex CLI is built around a single vendor's models and a coding workflow. kardolus/chatgpt-cli is built around a provider abstraction: the same binary talks to OpenAI, Azure, Perplexity, LLaMA, 302 AI and Atlas Cloud, and the --target flag switches between them. If you only ever use one provider, that abstraction is overhead you are paying for without collecting. If you switch providers, or need to keep a client working across several, it is the whole point. A second comparison is against shelling out to curl: curl gives you the raw request and no history management, thread isolation or sliding window, which means you rebuild all of that yourself.
Licence, maintenance and what an upgrade costs
The project is MIT licensed, which permits commercial and private use, modification and redistribution provided the licence text and copyright notice are kept. That is a permissive arrangement with few obligations, but it says nothing about the models you call: provider terms, data retention and per-token pricing are separate agreements you enter into independently of this repository. Nothing here is legal advice, and if you are embedding the binary in a product, the licence file in the repository root is the document to read.
Maintenance is active. The last push to the main branch was on 2026-09-01, and the most recent release in the repository is v1.11.0 from 2026-07-31, preceded by v1.10.12 on 2026-07-04 and v1.10.11 on 2026-03-22. That cadence is frequent enough that pinning a version is worth considering if you script against the binary.
The upgrade cost is concentrated in configuration. The project uses Viper for layered configuration and Cobra for the command surface, and the README documents a --target mechanism for switching between named configurations, plus a custom config and data directory option. That layering is convenient day to day, but it means a rename in a config key can break a setup silently, because the old key simply stops being read. The repository has a vendor directory and a Makefile with unit, integration and contract targets, so building and testing locally is a documented path rather than an adventure.
Editorial conclusion
Adopt it if you want one CLI, one YAML config and thread-scoped history across several providers, and if you are willing to read the agent policy before enabling agent mode. Skip it if you need a stable automation surface rather than an experimental one, or if you cannot supply an API key. Verify first that your provider's model names match the config keys, that the workdir sandbox covers the directories the agent will touch, and that your token budget is set before the first agent run.
Frequently asked questions
How do I install kardolus/chatgpt-cli on macOS?
The README documents a Homebrew formula for macOS. Direct download builds are also listed for Apple Silicon and macOS Intel chips, and the binary is invoked as chatgpt.
How do I install kardolus/chatgpt-cli on Windows?
The installation section lists a direct download for Windows amd64. There is no Homebrew path for Windows in the README.
Is kardolus/chatgpt-cli free?
The software itself is MIT licensed, so the binary is free to use and modify. The models it calls are not: the README's configuration sections all assume a provider API key, and provider pricing is separate from the licence.
What is kardolus/chatgpt-cli?
It is a multi-provider command-line interface for working with LLMs, written in Go. The README lists OpenAI, Azure, Perplexity and LLaMA among the supported providers, with streaming, interactive chat, prompt files, image and audio input, MCP tool calls and an experimental agent mode.
How do I set up kardolus/chatgpt-cli?
The README's Getting Started section covers the first run, and the configuration section documents general settings, per-provider settings and the agent configuration. The README states that custom context can be piped in from local files, standard input or another program.
How do I use kardolus/chatgpt-cli?
The README describes query mode for single input-output interactions, interactive mode for a conversational session, and streaming mode for real-time output. Model listing is available with the -l or --list-models flag.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kardolus-chatgpt-cli)