# openai-harmony: Rust Renderer and Parser for the gpt-oss Prompt Format

> openai-harmony is the reference Rust library, with Python bindings, for rendering and parsing the harmony response format that OpenAI's gpt-oss open-weight models were trained on. It is the correct tool for anyone building their own inference stack for gpt-oss outside of a hosted API or a pre-integrated inference server.

**openai/harmony** — Renderer for the harmony response format to be used with gpt-oss

- Repository: https://github.com/openai/harmony
- Stars: 4,509 · Forks: 311
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/openai-harmony

## What openai-harmony provides and who needs it

The gpt-oss models were trained on the harmony response format, which defines how conversations, reasoning channels, tool calls, and structured outputs are encoded as token sequences. The README is explicit about when this library is and is not needed: if you use the OpenAI Responses API or an inference provider such as HuggingFace, Ollama, or vLLM, the provider handles the formatting and this library is not required. If you are building your own inference stack directly, you need a correct implementation of the format.

openai-harmony provides that implementation in two ways: a Rust crate with the core rendering and parsing logic, and a Python package that wraps the Rust core through PyO3 bindings. The README describes the design goal as consistent formatting: a shared implementation for both rendering and parsing ensures that the token sequences produced match what the model was trained to expect, without round-trip loss. The README also advertises that the heavy lifting happens in Rust, giving Python users native-speed tokenization.

The Python package is published to PyPI as openai-harmony and includes typed stubs for IDE support. The Rust crate is not yet published to crates.io; it is consumed via a git dependency. The project is licensed under Apache-2.0 and the last push was on 2026-04-08.

## The harmony format: channels, roles, and tool namespaces

The README includes an example of the harmony format's wire representation. Conversations are delimited using special tokens:

```text
<|start|>system<|message|>You are ChatGPT, a large language model trained by OpenAI.
Knowledge cutoff: 2024-06
Current date: 2025-06-28

Reasoning: high

# Valid channels: analysis, commentary, final. Channel must be included for every message.
Calls to these tools must go to the commentary channel: 'functions'.<|end|>
```

The format supports multiple output channels in a single conversation: analysis, commentary, and final. Every message must include a channel assignment. Tool calls are routed to specific channels (commentary for function calls in the example). The format also supports a developer role with a separate instruction namespace, tool definitions with typed namespaces, and structured outputs. This multi-channel architecture is what enables the model to produce chain-of-thought reasoning separately from its final response, with each channel of output independently accessible.

The README notes that the format is designed to mimic the OpenAI Responses API, so teams already familiar with that API will recognize the conversation structure. The key difference is that the harmony format is the wire-level token representation rather than a JSON API payload: using gpt-oss without this format will cause the model to produce incorrect output.

## Installing and using openai-harmony in Python

The Python package installs from PyPI:

```bash
pip install openai-harmony
# or if you are using uv
uv pip install openai-harmony
```

The Python API mirrors the Rust data structures through PyO3 bindings. To render a conversation to tokens:

```python
from openai_harmony import (
    load_harmony_encoding,
    HarmonyEncodingName,
    Role,
    Message,
    Conversation,
    DeveloperContent,
    SystemContent,
)
enc = load_harmony_encoding(HarmonyEncodingName.HARMONY_GPT_OSS)
convo = Conversation.from_messages([
    Message.from_role_and_content(
        Role.SYSTEM,
        SystemContent.new(),
    ),
    Message.from_role_and_content(
        Role.DEVELOPER,
        DeveloperContent.new().with_instructions("Talk like a pirate!")
    ),
    Message.from_role_and_content(Role.USER, "Arrr, how be you?"),
])
tokens = enc.render_con
```

The Python package depends on pydantic>=2.11.7, as shown in pyproject.toml. Type stubs are included for static analysis. The Python and Rust test suites maintain 1-to-1 parity, meaning each test in tests.rs has an equivalent in the Python tests/ directory. This parity guarantee matters for downstream users who need to verify that switching from Python to Rust rendering produces identical token sequences. A demo subdirectory (demo/harmony-demo/) provides a working example application that can be run locally.

## Using openai-harmony in Rust

The Rust crate is not published to crates.io. The Cargo.toml dependency uses the git source:

```toml
[dependencies]
openai-harmony = { git = "https://github.com/openai/harmony" }
```

The Rust API provides the same HarmonyEncodingName, load_harmony_encoding, Conversation, Message, and Role types as the Python binding:

```rust
use openai_harmony::chat::{Message, Role, Conversation};
use openai_harmony::{HarmonyEncodingName, load_harmony_encoding};

fn main() -> anyhow::Result<()> {
    let enc = load_harmony_encoding(HarmonyEncodingName::HarmonyGptOss)?;
    let convo =
        Conversation::from_messages([Message::from_role_and_content(Role::User, "Hello there!")]);
    let tokens = enc.render_conversation_for_completion(&convo, Role::Assistant, None)?;
    println!("{:?}", tokens);
    Ok(())
}
```

The Rust core handles all tokenization work. The crate type is declared as both rlib and cdylib in Cargo.toml, which allows the same source to function as a Rust library for native use and as a compiled native extension (openai_harmony.*.so) for Python. The Rust implementation uses fancy-regex for complex tokenization patterns, rustc-hash for fast hashing, tiktoken-compatible logic, and reqwest with rustls for any network operations, avoiding an OpenSSL dependency on CI runners.

## Developer setup: building and testing locally

Developing openai-harmony requires the Rust stable toolchain, Python 3.8 or later, and maturin (the build tool for PyO3 projects). The setup steps from the README:

```bash
git clone https://github.com/openai/harmony.git
cd harmony
python -m venv .venv
source .venv/bin/activate
pip install maturin pytest mypy ruff
maturin develop --release
```

The maturin develop command compiles the Rust crate and installs the Python package in editable mode in the active virtualenv, similar to pip install -e for a pure Python project. To run both test suites and confirm implementation parity:

```bash
pytest && cargo test
```

Additional development commands from the README for type checking and formatting:

```bash
mypy harmony
ruff check .
cargo fmt --all
```

The run_checks.sh and test_python.sh scripts at the root automate the combined check workflow. A test-data/ directory provides fixtures for the test suites, and the Cargo.lock file pins the exact dependency versions to ensure reproducible builds.

## Repository architecture: Rust core and PyO3 bindings

The README provides the full source layout. The src/ directory contains the Rust crate:

```text
.
├── src/
│   ├── chat.rs           # High-level data-structures (Role, Message, …)
│   ├── encoding.rs       # Rendering & parsing implementation
│   ├── registry.rs       # Built-in encodings
│   ├── tests.rs          # Canonical Rust test-suite
│   └── py_module.rs      # PyO3 bindings ⇒ compiled as openai_harmony.*.so
│
├── python/openai_harmony/ # Pure-Python wrapper around the binding
│   └── __init__.py       # Dataclasses + helper API mirroring chat.rs
│
├── tests/                # Python test-suite (1-to-1 port of tests.rs)
└── Cargo.toml
```

The registry.rs module holds the built-in harmony encodings, which allows the library to ship with the correct encoding for gpt-oss without requiring a network call at runtime. The py_module.rs file is the PyO3 boundary: it translates between Rust types and Python objects. The pure-Python wrapper in python/openai_harmony/__init__.py provides the dataclass-based helper API that users interact with, keeping the Python surface clean and independent from the native extension details.

## Current limitations and what this library is not

The library is specific to the gpt-oss model series. It implements the harmony format for that model family and is not a general-purpose prompt formatting library for other models or for the broader OpenAI API surface. Teams that use GPT-4o, o-series models, or any other provider's models do not need this library. The README links to the OpenAI cookbook article on the harmony format for the full specification; openai-harmony is the implementation, not the specification document.

The Rust crate is not on crates.io, which means it cannot be pinned to a semver version through the standard Rust package registry. Projects that depend on a stable Rust API must manage the git dependency directly or vendor the source. There are no GitHub releases as of the last push on 2026-04-08, so version tracking requires pinning by commit hash in Cargo.toml rather than relying on published release tags.

The WASM binding target (wasm-bindgen, serde-wasm-bindgen, and wasm-bindgen-futures) appears as an optional Cargo feature set, and a javascript/ directory exists in the repository. The README does not document JavaScript usage in depth, and the npm package.json is not described in the available repository content. The Python package on PyPI is the primary non-Rust distribution channel at this stage. The demo/ directory contains a FastAPI-based example application using uvicorn, which illustrates how openai-harmony integrates into a Python web service that proxies requests to a gpt-oss model.

## How openai-harmony compares to manual format construction

Teams building custom gpt-oss inference stacks face a choice: implement the harmony format themselves from the specification, or use this reference library. Implementing it manually means parsing the <|start|>, <|message|>, and <|end|> token sequences correctly, maintaining a registry of the built-in encodings, and ensuring that rendering and parsing produce identical round-trip results. The README's argument for using this library is that it provides all of these as a tested, consistent implementation backed by the canonical Rust core.

A comparison point is how teams handle prompt formatting for other model families that do not have an official tokenization library. For those models, teams typically rely on the tiktoken library for tokenization and implement prompt templates manually. The gpt-oss series takes a different approach by publishing the rendering logic as an open-source library that is the reference for what the model expects. This reduces the risk of subtle formatting errors that would cause unexpected model behavior, such as incorrect channel assignments or malformed tool call preambles.

## Conclusion

openai-harmony is required only when running gpt-oss models outside of a provider that handles the format automatically. If you use the OpenAI Responses API, HuggingFace, Ollama, or vLLM, the README states you do not need this library. If you are building a custom inference server or a tokenization layer, openai-harmony is the reference implementation. The last push was on 2026-04-08 and there are no GitHub releases; the Rust crate is currently installed from source via the git dependency in Cargo.toml.

## FAQ

### When do I need openai-harmony instead of using the OpenAI Responses API?

The README states that if you use the OpenAI Responses API or an inference provider such as HuggingFace, Ollama, or vLLM, you do not need this library because the provider handles the formatting. You need openai-harmony only when building your own inference solution directly against gpt-oss model weights.

### Is openai-harmony available on crates.io?

The README's Rust installation example uses the git dependency form, which indicates the crate is not published to crates.io. There are no GitHub releases as of the last push on 2026-04-08, so Rust projects must depend on the library via a git reference in their Cargo.toml.

### Does openai-harmony work with Python versions below 3.8?

The pyproject.toml specifies requires-python = '>=3.8', and the PyO3 bindings are compiled with the abi3-py38 feature, which produces a single wheel compatible with Python 3.8 and all later versions. Python versions below 3.8 are not supported.

## Sources

- [Issues](https://github.com/openai/harmony/issues)
- [License: Apache-2.0](https://github.com/openai/harmony/blob/main/LICENSE)
- [openai/harmony on GitHub](https://github.com/openai/harmony)
- [README](https://github.com/openai/harmony/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/openai-harmony
