# llm-gemini: Google Gemini Models in the LLM CLI

> llm-gemini is a plugin for Simon Willison's LLM command-line tool that adds Google's Gemini model family to the llm ecosystem. It exposes Gemini's multimodal inputs, server-side tools such as code execution and Google Search grounding, and structured output, all through the same command-line interface used for other models in the llm plugin ecosystem.

**simonw/llm-gemini** — LLM plugin to access Google's Gemini family of models

- Repository: https://github.com/simonw/llm-gemini
- Stars: 462 · Forks: 57
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/simonw-llm-gemini

## What llm-gemini Provides and Who It Is For

The LLM CLI is a command-line tool and Python library by Simon Willison that provides a uniform interface for querying different AI models. It handles API key storage, response logging, prompt templating, and plugin-based model integration. The same llm command that queries one provider's models can query another provider's models after installing the relevant plugin, without changing the command structure. Response logging means every query and its output is saved locally, so a developer can retrieve previous outputs with `llm logs` without re-running the prompt.

llm-gemini adds Google's Gemini model family to this interface. The core use case is a developer who already uses llm for ad-hoc queries or shell scripting and wants to switch a specific task to a Gemini model without leaving the llm ecosystem or writing API integration code. The plugin also exposes Gemini-specific server-side capabilities that distinguish it from other llm plugins: code execution, Google Search grounding, and URL context are not available through the llm interface on other model providers.

The plugin is maintained by Simon Willison, the author of the LLM CLI. Version 0.34 was released on 2026-09-02, and the release history shows versions 0.32, 0.33, and 0.34 all released in 2026, indicating consistent maintenance. The Apache-2.0 license permits commercial and open-source use without restriction beyond attribution.

## Installing llm-gemini and Configuring the API Key

The plugin installs into the same Python environment as the llm tool, using the llm install mechanism:

```bash
llm install llm-gemini
```

Configuring the Gemini API key uses llm's built-in key storage:

```bash
llm keys set gemini
```

The key is pasted at the interactive prompt and stored in llm's key store. For non-interactive environments such as CI pipelines, the key can be supplied via the environment variable `LLM_GEMINI_KEY` instead. The plugin's pyproject.toml registers the plugin as an entry point under `[project.entry-points.llm]` with the identifier `gemini = "llm_gemini"`, which is how the llm CLI discovers the plugin after installation. This entry point mechanism means the plugin is automatically active after installation with no configuration step beyond setting the API key.

The plugin depends on llm version 0.32 or later, Python 3.10 or later, and the httpx HTTP client and ijson streaming JSON parser. These are installed automatically as dependencies. The minimum llm version of 0.32 is significant: anyone running an older llm installation will need to upgrade it before installing llm-gemini, which may affect other plugins in the same environment that have different version constraints.

## Available Models and Running the First Query

The plugin exposes models in the Gemini Flash, Gemini Flash Lite, Gemini Pro preview, and Gemma instruction-tuned families. Model identifiers use the `gemini/` prefix, such as `gemini/gemini-flash-latest`, `gemini/gemini-2.5-flash`, `gemini/gemma-4-31b-it`, and `gemini/gemini-3.8-flash`. Each model also has an alias without the prefix, so `gemini-flash-latest` and `gemini/gemini-flash-latest` refer to the same model. The Gemma models (gemma-4-31b-it, gemma-4-26b-a4b-it) are Google's open-weight models accessible through the same Gemini API endpoint, which means they are available through the plugin without any additional configuration.

A basic query uses the `-m` flag to specify the model:

```bash
llm -m gemini-flash-latest "A short joke about a pelican and a walrus"
```

To avoid specifying the model on every call, set a default:

```bash
llm models default gemini-flash-latest
```

After that, `llm "A joke about a pelican and a walrus"` uses the chosen model without the `-m` flag. The model list in the README is generated from the plugin's own code using the cogapp tool, so it reflects the models the plugin actually registers rather than a manually maintained list. The `gemini-flash-latest` alias tracks Google's current recommended Flash model, which means scripts that use it will automatically benefit from Google's model updates without requiring a version bump in the command. Versioned identifiers like `gemini/gemini-2.5-flash` pin to a specific release and will fail when Google removes that version from the API.

## Multimodal Input: Images, Audio, Video, and YouTube

Gemini models accept non-text inputs through llm's `-a` attachment flag. An image from a local file:

```bash
llm -m gemini-flash-latest 'extract text' -a image.jpg
```

An image from a URL:

```bash
llm -m gemini-flash-latest 'describe image' \
  -a https://static.simonwillison.net/static/2024/pelicans.jpg
```

Audio transcription:

```bash
llm -m gemini-flash-latest 'transcribe audio' -a audio.mp3
```

Video description:

```bash
llm -m gemini-flash-latest 'describe what happens' -a video.mp4
```

YouTube video URLs work as attachments as well, which is a Gemini-specific capability the API supports natively. YouTube videos are processed at low media resolution by default. The `-o media_resolution X` option changes this to `medium`, `high`, or `unspecified`. The README links to a sample output demonstrating transcript and summary extraction from a YouTube video, showing that the plugin passes the URL directly to the Gemini API's native YouTube handling. The Gemini prompting guide linked from the README includes detailed advice on multi-modal prompting strategies that apply to the `-a` flag regardless of the input type.

## Server-Side Tools: Code Execution, Google Search, and URL Context

Three Gemini server-side tools are enabled via the `-T` flag. Code execution lets the model write Python code, run it in a secure sandbox, and use the result in its response:

```bash
llm -m gemini-3.6-flash -T CodeExecution \
'use python to calculate (factorial of 13) * 3'
```

Google Search grounding lets the model query Google and use the search results when formulating its answer:

```bash
llm -m gemini-3.6-flash -T GoogleSearch \
  'What happened in Ireland today?'
```

The plugin retains Gemini's raw `groundingMetadata` on the response. The README notes that the full metadata, including details about grounded results, can be inspected by running `llm logs -c --json` after a grounded query. This metadata includes additional information about which search results were used and how they were incorporated. Using the Google Search tool may carry additional usage requirements under Google's documentation, which the README directs users to review.

URL context is a third server-side tool that allows the model to fetch and use content from URLs during execution. The URL context tool is enabled with `-T URLContext`. These server-side tools are invoked on Google's infrastructure, not locally, and their availability may depend on which Gemini model is used. Not all models in the family support all three tools: checking the Gemini API documentation for a specific model before scripting around these tools is advisable.

## Structured Output and JSON Schema

The plugin supports two approaches for structured output. The `-o json_object 1` option forces the response to be valid JSON:

```bash
llm -m gemini-flash-latest -o json_object 1 \
  '3 largest cities in California, list of {"name": "..."}'
```

The README shows the output for that command:

```json
{"cities": [{"name": "Los Angeles"}, {"name": "San Diego"}, {"name": "San Jose"}]}
```

The `--schema` flag constrains output to a specific field structure:

```bash
llm -m gemini-flash-latest --schema 'name,age int,bio' 'invent a dog'
```

The schema syntax is compact: a comma-separated list of field names where type suffixes such as `int` are optional. Both structured output modes are part of the llm CLI framework, and llm-gemini exposes them for Gemini models using the same interface used by other model plugins in the ecosystem. Schema-constrained output is useful when feeding llm responses into downstream tools that expect a fixed data structure, such as a shell pipeline that reads a specific JSON field with jq.

## Limitations and How llm-gemini Compares to the Gemini SDK

llm-gemini is not a standalone library. It depends entirely on the llm CLI being installed and contributes no Python API of its own. Developers who need to integrate Gemini into an application, call Gemini from a web server, or use streaming responses in production code should use Google's official Gemini Python SDK, which provides the full API surface including session management, token counting, and function calling. The llm-gemini plugin exposes no direct equivalent for those features.

Google's Gemini SDK is the first-party Python library for the Gemini API. It is designed for embedding in applications and gives direct access to all API parameters. llm-gemini is designed for interactive command-line use and shell scripting, where the llm ecosystem's key management, response logging, and unified model interface add value over raw SDK calls. A developer who needs to compare a Gemini response with a Claude response on the same prompt can do so with two llm commands rather than setting up two separate API integrations.

The ijson dependency listed in pyproject.toml is a streaming JSON parser, which suggests the plugin handles Gemini's streaming response format incrementally rather than buffering the full response before printing. This matters for long-form outputs where waiting for a complete buffer before showing any output would create noticeable delay in interactive use. The httpx dependency handles the HTTP transport layer for API calls. Both are installed automatically when the plugin is installed and do not require separate configuration steps from the user.

## Conclusion

llm-gemini is the right choice for developers already using the llm CLI who want to route Gemini queries through the same interface, key management, and logging they use for other models. It is not a standalone SDK: it requires llm version 0.32 or later and provides no Python API for embedding in applications. Before adopting it, verify your llm environment is at 0.32+, and check that the specific model identifier you plan to use is still active in Google's API, since model names change as Google releases new versions. For application-level Gemini integration, the official Gemini SDK is the appropriate starting point.

## FAQ

### How do you use llm-gemini to run a Gemini model?

Install the plugin with `llm install llm-gemini`, set your API key with `llm keys set gemini`, then run `llm -m gemini-flash-latest "your prompt"`. Set a default model with `llm models default gemini-flash-latest` to avoid specifying `-m` on every call.

### What is llm-gemini?

llm-gemini is a plugin for Simon Willison's LLM command-line tool that adds Google's Gemini model family. It supports multimodal inputs (images, audio, video, YouTube), server-side tools including code execution and Google Search grounding, and structured JSON output, all through the same llm command interface.

### Is llm-gemini free to use?

The plugin is open source under the Apache-2.0 license. Using it requires a Google Gemini API key, obtainable from Google AI Studio. Google's API has its own pricing and free-tier limits set by Google, separate from the plugin.

## Sources

- [Issues](https://github.com/simonw/llm-gemini/issues)
- [License: Apache-2.0](https://github.com/simonw/llm-gemini/blob/main/LICENSE)
- [README](https://github.com/simonw/llm-gemini/blob/main/README.md)
- [Releases](https://github.com/simonw/llm-gemini/releases)
- [simonw/llm-gemini on GitHub](https://github.com/simonw/llm-gemini)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/simonw-llm-gemini
