Model or dataset
leeguooooo/chatgpt-imagegen avatar
leeguooooo/chatgpt-imagegen

chatgpt-imagegen: generate images from the command line with your ChatGPT subscription

Use your ChatGPT subscription to generate images from the command line — no OPENAI_API_KEY, no gateway, no daemon. Zero-dep Python CLI + AI-agent skill.

366 stars43 forksPythonMIT

At a glance

What is it?
A zero-dependency Python CLI and AI-agent skill that drives your logged-in ChatGPT session instead of an API key. Here is how the backends differ, what the codex-only flags actually do, and where the tool stops being the right choice.
Who is it for?
Adopt it if you already pay for ChatGPT or Codex and want image generation inside a shell script or an agent workflow without provisioning an API key. Skip it if you need a stable, versioned HTTP contract, because the web backend depends on a logged-in Chrome session that OpenAI can change at any time, and the README does not document rollback for a bad upgrade.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 14 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The API-key tax chatgpt-imagegen removes

Most image-generation tooling assumes you hold an OpenAI API key and pay per call. chatgpt-imagegen assumes the opposite: you already have a ChatGPT subscription, possibly the free tier, and you want to spend that instead. The README states the default backend drives the normal ChatGPT web chat, "where even free-tier users get image generation." That is the whole pitch, and it is narrower than it sounds. This is a tool for people who already sit inside a terminal or an agent session and want a PNG to appear at a path, not for teams building a product on top of an image endpoint. The README's own example is a single command that writes cat.png and prints the byte count, the size and the quality the backend reported. If your workflow is "I need 400 images a day with a retry policy and a cost dashboard," this is the wrong layer.

How the web, codex, gemini and agy backends differ

The CLI is one Python file using only the standard library, but it is not self-contained in the sense that matters: it needs a backend. The default, web, drives your logged-in Chrome and, per the README, spends no Codex usage. The codex backend is the headless fallback and does consume the metered Codex bucket. Two optional backends, gemini and agy, drive gemini.google.com in Chrome or the Antigravity CLI respectively, and the README notes they bill separate quotas from each other, so one can cover for the other. Neither Gemini backend is ever selected automatically; you ask for it by name. That is a deliberate design choice and a reasonable one, because silently spending a different vendor's quota would be worse than making the user type a flag. The practical consequence is that your output is not deterministic across machines. A colleague on codex and you on web can get different pixel dimensions from the same prompt, because the surfaces expose different controls.

Installing chatgpt-imagegen and generating a first image

The README gives two install paths. For agents, a single npx command drops the skill into Claude Code, Codex, Cursor and similar tools.

bash
npx skills add leeguooooo/chatgpt-imagegen -g

After that you ask in natural language and the skill handles the rest. For a standalone CLI there is no pip and no virtualenv: clone the repository and install the single script onto your PATH. You need Python 3.10 or newer.

bash
git clone https://github.com/leeguooooo/chatgpt-imagegen
sudo install chatgpt-imagegen/chatgpt-imagegen /usr/local/bin/chatgpt-imagegen

Before generating anything, check what is actually ready. The doctor subcommand reports the backend state and, separately, whether the animation dependencies are present.

bash
chatgpt-imagegen doctor

A first real run is a prompt plus an output path. The README's example prints a saved line with the byte count, the size and the quality the backend used.

bash
chatgpt-imagegen "a watercolor cat sitting on a windowsill" -o cat.png

To capture the path instead of reading it off the screen, use --quiet with command substitution, which the README shows as OUT=$(chatgpt-imagegen "icon" --quiet).

The GPT Image 2.5 flags are requests, not guarantees

The most easily misread part of the README is the set of codex-only knobs: --image-model, --quality, --background transparent, --compression, --action and --partial-images. They are optional and codex-only, since the web and Gemini surfaces have no such controls. More importantly, the README says the codex backend rewrites them server-side, observing gpt-image-2-codex and auto, so "they rarely survive there." The saved line prints the model, quality and size the backend actually used, which is the only honest way to know what you got. On codex, the flag that does matter is --model, the driver model that calls the tool. It defaults to gpt-5.6-luna, described as a fast and affordable Codex model, on the reasoning that a frontier coding model just burns the metered Codex bucket. Every run prints the tokens it cost. If you are budgeting, read that line rather than assuming the flags you passed were honored.

Styles, characters and the animate subcommand

Styles are the feature that separates this from a thin wrapper. A style can pin a character, not just a look: the README shows style add pip --kind character --ref pip-ref.png, after which the same fox returns in a new scene. Styles can also be pulled from a public gallery at drawstyle.leeguoo.com without updating the script, via --style-online, and searched with style search. The animate subcommand is more constrained than it first appears. It asks the image model for a strict 4x2 sprite sheet, crops eight equal frames, rejects obvious subject drift and writes a ping-pong loop, defaulting to animated WebP. Post-processing needs ImageMagick's magick binary, and WebP output additionally needs img2webp from libwebp. The source sprite PNG is always kept beside the animation, which is a sensible decision for debugging a bad loop. The README does not describe what happens when frame rejection fires on most frames, so treat animate as the least predictable part of the tool.

Where chatgpt-imagegen breaks, and what to use instead

The failure mode is structural. The web backend depends on driving a logged-in Chrome session against the ChatGPT web UI, and the README does not document rollback if an upgrade goes wrong. It does document one real upgrade trap: on 0.23.1 or earlier the self-update only looked for a global skills binary and gave up when it was missing, so those versions cannot deliver their own fix and need a one-time bootstrap with npx -y skills update chatgpt-imagegen. That is a concrete, verifiable example of the risk in depending on a moving surface. If you need a contract you can version, the OpenAI Images API is the alternative, and the difference is not cosmetic: it is a documented HTTP endpoint with an API key, per-call billing and no browser session, which means it survives UI changes but costs money per image and requires key management. A second alternative is a self-hosted diffusion stack, which gives you full control over the model and no vendor dependency, at the cost of GPU hardware and prompt-tuning work. chatgpt-imagegen wins only when the subscription is already paid for and the browser session is something you can keep alive.

Licence, maintenance and the cost of staying current

The project is MIT licensed, so you can vendor the single script into your own repository and modify it; keep the licence text with it. This is not legal advice, and the licence covers the code, not your ChatGPT or Gemini account terms, which govern whether automated use of those web surfaces is permitted at all. On maintenance, the last push was on 2026-09-11, and releases v0.25.0 through v0.27.0 all landed that day, so the project is moving. The upgrade cost is real but small: chatgpt-imagegen update runs the skills manager for you, using npx when skills is not on PATH, and interactive runs check for a newer version at most once a day and upgrade automatically for the next run. Two environment variables control this. CHATGPT_IMAGEGEN_NO_AUTO_UPDATE=1 disables installation but keeps the check and notice; CHATGPT_IMAGEGEN_NO_UPDATE_CHECK=1 disables both. --quiet and --no-progress never upgrade in the background, which is what you want in CI.

Editorial conclusion

Adopt it if you already pay for ChatGPT or Codex and want image generation inside a shell script or an agent workflow without provisioning an API key. Skip it if you need a stable, versioned HTTP contract, because the web backend depends on a logged-in Chrome session that OpenAI can change at any time, and the README does not document rollback for a bad upgrade. Before relying on it, run chatgpt-imagegen doctor to confirm your backend and, if you plan to use animate, that magick and img2webp are on PATH.

Frequently asked questions

How do I use chatgpt-imagegen without an OpenAI API key?

The tool is built around that constraint: the README states it generates images with your ChatGPT subscription and needs no OPENAI_API_KEY. The default web backend drives your logged-in Chrome session, and the README notes even free-tier ChatGPT users get image generation there.

What is the chatgpt-imagegen model used for image generation?

It depends on the backend. On codex you can pass --image-model with values such as sunburst or flare, but the README says the backend rewrites these server-side, so they rarely survive; the saved line prints the model the backend actually used. On codex the --model flag selects the driver model, defaulting to gpt-5.6-luna.

How does chatgpt-imagegen compare with using Gemini for image generation?

chatgpt-imagegen supports both: --backend gemini drives a logged-in gemini.google.com Chrome profile and --backend agy uses the Antigravity CLI. The README says the two Gemini backends bill separate quotas from each other and neither is chosen automatically. It also notes Gemini text-to-image output carries a visible bottom-right watermark, while image-to-image does not.

What is chatgpt-imagegen called?

The project name is chatgpt-imagegen, and the repository is leeguooooo/chatgpt-imagegen. The README describes it as a zero-dependency Python CLI plus an AI-agent skill for generating images with a ChatGPT subscription.

Official sources

  1. Issues
  2. leeguooooo/chatgpt-imagegen on GitHub
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/leeguooooo-chatgpt-imagegen.svg)](https://hysenlabs.com/projects/leeguooooo-chatgpt-imagegen)