# ima2-gen: Local-First Visual Generation Studio for Multiple AI Providers

> ima2-gen is an npm package and Mac app that runs a local server for AI image and video generation, connecting to OpenAI, Grok, Gemini, NovelAI, and ComfyUI from a single interface. It keeps generated images on your machine and tracks every prompt, timing, and provider choice so work can be resumed or branched without starting over.

**lidge-jun/ima2-gen** — Local-first visual generation runtime and studio for people and coding agents, with reproducible image and video workflows across multiple providers.

- Repository: https://github.com/lidge-jun/ima2-gen
- Website: https://lidge-jun.github.io/ima2-gen/
- Stars: 848 · Forks: 135
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/lidge-jun-ima2-gen

## A Local Server That Connects to Multiple Image and Video Providers

ima2-gen runs a small HTTP server on your machine and acts as a local hub for visual generation across several provider accounts. It connects to OpenAI OAuth and API, Grok OAuth and API, Antigravity CLI, Gemini API, AtlasCloud, MiniMax, NovelAI, and registered ComfyUI workflows. Runway and Higgsfield are available through separate MCP-backed integrations rather than the built-in provider list.

The key design principle is that prompts and reference images go only to the provider you pick for each job. No central service aggregates your creative work. Every generated image is stored at `~/.ima2/generated`, and each result keeps its full metadata: the prompt, the provider and model used, timing, and settings. This makes it possible to copy a result, continue from it, or animate it without reconstructing the original parameters from memory.

This approach is most useful when you routinely work with more than one provider and want to compare results or mix providers across stages of a project. A team or solo creator using only one provider gains less from the multi-provider routing.

## How Providers, Lanes, and Models Are Organized

ima2-gen uses a two-part naming convention for models: a lane identifier followed by a model slug, separated by a slash. The README gives `oauth/gpt-6-luna` as an example. Lanes map to provider access methods (OAuth vs. API key), and models within a lane correspond to the offerings from that provider. Running `ima2 models` lists what is available with your current configuration.

Defaults are set per output type. To set a default image model:

```bash
ima2 defaults set image oauth/gpt-6-luna
```

To set a default video model:

```bash
ima2 defaults set video grok/grok-imagine-video-1.5
```

Once defaults are configured, `ima2 gen` and `ima2 video` work without further flags. Without a configured default, `ima2 gen` fails closed with a `NO_DEFAULT_MODEL` error rather than silently picking a provider. This behavior prevents unexpected API billing on a provider you did not choose. Any call can override the default by passing `--model <lane>/<model>` or `--provider <lane>` explicitly.

The setup wizard, invoked through `ima2 setup`, offers four authentication paths: GPT OAuth, Grok OAuth, both, or a web-based setup. Video generation specifically requires Grok OAuth, so users who configured only GPT OAuth need to run `ima2 grok login` separately to add video capability.

## Installing ima2-gen: npm, Mac App, and Platform Scripts

The standard npm installation installs the package globally and then runs setup and the server:

```bash
npm install -g ima2-gen
ima2 setup
ima2 serve
```

The server binds to port 3333 by default. If that port is taken, ima2-gen finds the next free port and writes the real URL to `~/.ima2/server.json`. Running `ima2 open` reads that file and opens the correct URL without guessing.

For macOS on Apple Silicon, a signed and Apple-notarized desktop app is available as a DMG download from the ima2 Desktop releases page. The Mac app includes its own runtime, so Node.js is not required on the host. On Intel Mac, Windows, and Linux, the npm path or a one-line installer is the recommended approach.

One-line installers check the Node.js version floor, install Node LTS if needed, and then install ima2-gen:

```bash
curl -fsSL https://lidge-ai.github.io/ima2-gen/install-mac.sh | bash
```

For Linux or WSL:

```bash
curl -fsSL https://lidge-ai.github.io/ima2-gen/install-linux.sh | bash
```

For Windows via PowerShell:

```powershell
irm https://lidge-ai.github.io/ima2-gen/install-windows.ps1 | iex
```

To update, stop the server with Ctrl+C or `ima2 stop` from another terminal, then install the latest version:

```bash
npm install -g ima2-gen@latest
```

The README notes that Ctrl+C performs a clean shutdown: it closes the database, stops child processes, and releases file locks.

## Node Graph for Branching, Canvas Mode for Cleanup

ima2-gen provides two visual tools beyond simple prompt-and-result generation. The node graph lets you take a finished image and push it in several directions at once. Each branch keeps its parent image as the source, and branches remember their relationship so nothing gets overwritten when you explore variations. Finished jobs match back to the graph by request ID, which means reloads and version conflicts do not lose results. The graph supports up to 500 nodes and 1000 edges, as documented in the `.env.example` configuration.

Canvas Mode targets image cleanup tasks: zoom, pan, annotation with hover highlighting, erase, group operations, undo, sticky notes, background removal, and export with preserved alpha transparency or a matte color. Export formats include SVG with an embedded raster layer and a traced vector SVG that flattens the composition into real vector paths. The README notes that transparency from GPT providers is verified server-side before the app reports a result as transparent, which guards against provider API inconsistencies.

Multimode allows generating several candidates from one prompt simultaneously and watching them fill slots as they arrive. The Prompt Studio manual documents every control, multimode recipes, direct mode, and reasoning effort settings.

## Running ima2-gen in Docker

A Dockerfile is included in the repository. Build and run it with:

```bash
docker build -t ima2-gen .
docker run -d -p 3333:3333 -e IMA2_LAN_TOKEN=change-me -v ima2-data:/data ima2-gen
```

The `IMA2_LAN_TOKEN` environment variable is required. The server refuses to bind to a non-loopback host without it, which prevents accidental exposure on a network interface. The `.env.example` file in the repository root documents all available environment variables: `IMA2_PORT`, `IMA2_HOST`, `IMA2_CONFIG_DIR`, `IMA2_GENERATED_DIR`, `IMA2_MAX_PARALLEL`, and others.

The Docker image uses Node 22 on Debian Bookworm and runs `node server.js` directly rather than `ima2 serve`, because the `serve` subcommand enters an interactive setup wizard when no provider is configured, which would hang a container. A Docker Compose file is also included for compose-based deployments. Provider API keys can be passed as environment variables or configured through the web UI after the container starts.

## What ima2-gen Does Not Handle: Provider Limits and Missing Features

ima2-gen is a local client layer, not a provider. If a provider's API is unavailable or your account reaches a quota limit, ima2-gen cannot bypass those constraints. The tool also does not aggregate billing: each provider charges your account directly, and ima2-gen does not report consolidated cost figures across providers.

Reference images are limited by provider and by the client layer. For images, up to five references are supported. For video, up to fourteen references are allowed. Large images are compressed before upload, but the maximum decoded size for a reference is approximately 5.2MB as configured in the default environment variables.

Runway and Higgsfield integrations are described as separate MCP-backed integrations rather than first-class built-in providers. This means they require a running MCP setup rather than a direct API key configured in `ima2 setup`.

A real alternative is ComfyUI in standalone mode. ComfyUI runs locally and provides a node-based workflow editor for image generation without external API calls or per-image cloud costs. The difference is that ComfyUI focuses on ComfyUI-native workflows and local model weights, while ima2-gen focuses on routing to cloud providers with a unified interface and multi-provider comparison. ima2-gen can also use ComfyUI as one of its registered providers, but that is a narrower integration than running ComfyUI directly for its full workflow system.

## Maintenance, Versioning, and License

The last push to the repository was on 2026-09-27. The most recent releases include v3.23.1 and Desktop 3.23.2, both published on 2026-09-27. The package version in `package.json` is 3.23.2 with `npm@11.18.0` as the package manager. The version scheme suggests active release cadence.

The repository is structured as a TypeScript project with separate build steps for the server, CLI, and UI. The `bin/ima2.js` entry point is the CLI binary. The frontend lives in `ui/` and is built separately from the server. The `routes/`, `lib/`, `integrations/`, and `skills/` directories hold the server-side logic.

The license is MIT, which permits use, modification, and distribution without restriction. The repository includes a CHANGELOG.md, a SECURITY.md, and a CONTRIBUTING.md. The `AGENTS.md` file at the root documents how coding agents should interact with the project, reflecting the README's positioning of ima2-gen as a tool designed to be usable both by people and by coding agents.

## Conclusion

ima2-gen suits developers and creative professionals who work with multiple image generation providers and want a single interface that keeps every result reproducible and branchable on their own hardware. It is a wrong fit if you need a hosted service with no local server dependency, or if your primary provider is not one of the supported ones. Before starting, verify which OAuth or API credentials you have: video generation through `ima2 video` requires Grok OAuth, and generating images through the CLI requires running `ima2 defaults set` first or passing `--model` explicitly, otherwise the command fails closed with `NO_DEFAULT_MODEL`.

## FAQ

### What providers does ima2-gen support for image generation?

ima2-gen connects to OpenAI OAuth and API, Grok OAuth and API, Antigravity CLI, Gemini API, AtlasCloud, MiniMax, NovelAI, and registered ComfyUI workflows. Runway and Higgsfield are available through separate MCP-backed integrations.

### How does ima2-gen handle video generation, and what is required to enable it?

Video generation uses the `ima2 video` command and requires Grok OAuth authentication. If you configured only GPT OAuth during setup, run `ima2 grok login` to add video capability. The command fails closed with `NO_DEFAULT_MODEL` until a video default is set via `ima2 defaults set video`.

### Can ima2-gen run in a Docker container without an interactive setup wizard?

Yes, by running the container with `node server.js` directly instead of `ima2 serve`. The `ima2 serve` command enters an interactive wizard when no provider is configured, which would hang a container. The included Dockerfile and docker-compose.yml use the direct server entry point and require `IMA2_LAN_TOKEN` to be set in the environment.

## Sources

- [License: MIT](https://github.com/lidge-jun/ima2-gen/blob/main/LICENSE)
- [lidge-jun/ima2-gen on GitHub](https://github.com/lidge-jun/ima2-gen)
- [Project website](https://lidge-jun.github.io/ima2-gen/)
- [README](https://github.com/lidge-jun/ima2-gen/blob/main/README.md)
- [Releases](https://github.com/lidge-jun/ima2-gen/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lidge-jun-ima2-gen
