# openai/gpt-oss: OpenAI's Open-Weight Models and the Repository That Ships Them

> gpt-oss-120b and gpt-oss-20b are Apache-2.0 licensed open-weight models trained on the harmony response format. The openai/gpt-oss repository holds reference inference implementations, not the weights themselves.

**openai/gpt-oss** — gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI

- Repository: https://github.com/openai/gpt-oss
- Website: https://openai.com/open-models
- Stars: 20,435 · Forks: 2,161
- Language: Python
- License: Apache-2.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/openai-gpt-oss

## What gpt-oss actually is, and who it is for

The openai/gpt-oss repository is not where the model weights live. According to the README, the models are downloaded from Hugging Face, and the repository itself is described in pyproject.toml as "A collection of reference inference implementations for gpt-oss by OpenAI." That distinction matters when you evaluate it: you are looking at tooling, not a model artifact.

Two models are released. gpt-oss-120b is positioned for production, general purpose, high reasoning use cases, with 117B parameters and 5.1B active parameters, sized to fit a single 80GB GPU such as an NVIDIA H100 or AMD MI300X. gpt-oss-20b is positioned for lower latency and local or specialized use, with 21B parameters and 3.6B active parameters, and the README states it runs within 16GB of memory. Both are post-trained with MXFP4 quantization of the MoE weights, and all evals were performed at that quantization.

The audience is engineers who want open weights under a permissive license. The README lists fine-tuning, configurable reasoning effort, full chain-of-thought access, and agentic capabilities (function calling, web browsing, Python code execution, structured outputs) as the intended uses. If you need a hosted endpoint rather than weights you can inspect and modify, this is the wrong starting point.

## The harmony response format is a hard dependency, not a suggestion

The README is explicit: both models were trained using the harmony response format and "should only be used with this format; otherwise, they will not work correctly." This is the single most consequential constraint in the project, and it shapes every integration path.

Harmony is a conversation rendering layer. It takes system, developer, and user messages, renders them into token sequences for the model, and parses completion tokens back into structured entries such as assistant messages and tool calls. The openai-harmony package is a required dependency in pyproject.toml, so any Python install pulls it in. The README notes that if you use Transformers' chat template, harmony is applied automatically; if you call model.generate directly, you must apply it manually or use openai-harmony.

That design has a cost. Prompt templates you already have for other models will not transfer. Any pipeline that concatenates raw strings and ships them to the model is likely to produce degraded or malformed output. The benefit is that reasoning traces, tool calls, and final answers arrive as distinguishable structured entries rather than one undifferentiated string, which is what makes the documented agentic features tractable.

## Installing gpt-oss and running gpt-oss-20b through vLLM

The package requires Python 3.12 or newer. For a server-style deployment, the README points at vLLM and recommends uv for dependency management. The install command pins a gpt-oss-specific vLLM build and adds two extra indexes:

```bash
uv pip install --pre vllm==0.10.1+gptoss \
    --extra-index-url https://wheels.vllm.ai/gpt-oss/ \
    --extra-index-url https://download.pytorch.org/whl/nightly/cu128 \
    --index-strategy unsafe-best-match
```

After that, serving the smaller model is a single command. The README states it downloads the model automatically and starts an OpenAI-compatible web server:

```bash
vllm serve openai/gpt-oss-20b
```

For offline use, the README's example adds openai-harmony and renders the conversation before generation. The snippet below sets an environment variable to disable the FlashInfer sampler, loads the harmony encoding, builds a conversation with a developer instruction, and renders it for completion:

```python
import os
os.environ["VLLM_USE_FLASHINFER_SAMPLER"] = "0"

from openai_harmony import (
    HarmonyEncodingName,
    load_harmony_encoding,
    Conversation,
    Message,
    Role,
    SystemContent,
    DeveloperContent,
)

encoding = load_harmony_encoding(HarmonyEncodingName.HARMONY_GPT_OSS)
convo = Conversation.from_messages([
    Message.from_role_and_content(Role.SYSTEM, SystemContent.new()),
    Message.from_role_and_content(
        Role.DEVELOPER,
        DeveloperContent.new().with_instructions("Always respond in riddles"),
    ),
    Message.from_role_and_content(Role.USER, "What is the weather like in SF?"),
])
prefill_ids = encoding.render_conversation_for_completion(convo, Role.ASSISTANT)
```

The same example passes stop_token_ids from encoding.stop_tokens_for_assistant_actions() into SamplingParams so those tokens are excluded from the output, then calls encoding.parse_messages_from_completion_tokens on the returned token IDs to recover structured messages. What you should see is a sequence of JSON-serialized conversation entries rather than a plain string.

## The reference PyTorch, Triton and Metal implementations are not production paths

The repository ships three additional implementations, and the README is unusually direct about their status: they are "largely reference implementations for educational purposes and are not expected to be run in production." The optional dependency groups in pyproject.toml match that framing. triton pulls triton>=3.4, safetensors>=0.5.3 and torch>=2.7.0; torch pulls safetensors and torch; metal pulls numpy, tqdm, safetensors and torch.

This is a real limitation rather than a disclaimer. If you need throughput, batching, or a supported serving stack, the documented paths are vLLM or another listed client. The value of the reference code is that it shows how the architecture and harmony rendering fit together, which is useful if you intend to port the model to another runtime or verify that an implementation matches. Treating the Triton single-GPU path as a deployment target means owning performance work the project has not committed to.

One detail worth noting for anyone building from source: the build system is not plain setuptools. The build-system block points at a custom backend in _build, and tool.scikit-build drives a CMake build from the root CMakeLists.txt with -DGPTOSS_BUILD_PYTHON=ON and a Release build type. That is a heavier build than a pure-Python package, and it is the kind of thing that surfaces platform-specific problems early.

## Where gpt-oss is the wrong choice

The most common mismatch is treating gpt-oss as a self-hosted substitute for OpenAI's hosted API. The models are open-weight, but the README does not describe them as equivalent to any hosted model, and the harmony requirement means existing prompts and tool schemas need rework. If your application is built around a different message format, migrating is a project, not a config change.

Hardware is the second boundary. gpt-oss-120b needs a single 80GB GPU. That is a specific class of accelerator, not a general-purpose server. gpt-oss-20b fits in 16GB, which the README frames as consumer hardware territory and pairs with an Ollama path. If your target machine has less than that, neither model is the right fit at the stated quantization.

There is also a reasoning-effort trade-off. The README lists configurable reasoning effort at low, medium and high, tied to latency needs. Lower effort means less computation spent on the reasoning path, which is the point, but it also means the chain-of-thought you get for debugging is thinner. Teams that want full traces for auditing should plan around the higher settings and their latency.

Finally, the README states plainly that full chain-of-thought "is not intended to be shown to end users." If your product depends on displaying reasoning to users, that is a design decision the project pushes against, and you should decide deliberately rather than by default.

## gpt-oss versus Gemma and other open-weight families

The nearest comparison point people search for is gpt-oss versus Gemma, and the difference that matters is not benchmark scores but the interface contract. gpt-oss is trained on harmony and the README says it will not work correctly without it. That gives you structured reasoning and tool-call entries out of the box, at the cost of a format you must adopt wholesale.

A family that uses a conventional chat template lets you swap models behind an existing prompt pipeline with minimal change. With gpt-oss, the chat template route exists through Transformers, but the moment you move to raw generation you are responsible for rendering and parsing harmony yourself. That is the trade: more structure and more integration work.

The second difference is deployment shape. gpt-oss ships two sizes with stated memory targets, one of which is explicitly aimed at a single 80GB GPU. Families with a wider size ladder let you pick a model that matches whatever hardware you already have. With gpt-oss you choose between 16GB and 80GB, and there is no middle option documented in this repository.

Licensing is the third axis. gpt-oss is Apache-2.0, which the README highlights as permissive with no copyleft restrictions or patent risk. If your legal review has already cleared Apache-2.0 dependencies, this fits that pattern. Note that the repository also contains a USAGE_POLICY file; the README does not explain how it interacts with the license, so read it before shipping.

## Maintenance, licensing and what to verify before adopting

The repository is not archived, and the most recent push recorded is 2026-07-24. The latest release is v0.0.9 from 2026-01-13, following v0.0.8 in September 2025 and v0.0.7 two weeks earlier. The version number in pyproject.toml matches v0.0.9. A 0.0.x version series signals that interfaces can move, and the release cadence is not uniform, so pin your dependency rather than tracking the default branch.

Upgrade cost concentrates in two places. First, the vLLM install pins a specific build (vllm==0.10.1+gptoss) against nightly PyTorch wheels from a CUDA 12.8 index. That combination will need revisiting as those indexes move. Second, harmony is a required dependency, so changes to openai-harmony propagate into every integration. The offline example also sets VLLM_USE_FLASHINFER_SAMPLER=0, which is the kind of environment-level workaround that tends to be version-sensitive.

On licensing: the project is Apache-2.0, which permits commercial use and modification without copyleft obligations. That is the license file's role, not legal advice, and the presence of a separate USAGE_POLICY file in the repository root means your review should cover both documents rather than the LICENSE alone.

Before adopting, verify three things against your own environment: that your GPU matches the stated memory target for the model you pick, that your prompt and tool-call pipeline can be rebuilt around harmony, and that the pinned vLLM and PyTorch indexes resolve on your platform. The repository also includes a compatibility-test directory and a tests directory, which is where to look for the project's own expectations about supported configurations.

## Conclusion

Adopt gpt-oss when you need an Apache-2.0 open-weight model you can run on your own hardware and fine-tune, and when you can commit to the harmony response format. Do not adopt it if you expect a drop-in OpenAI API replacement or need models outside the two released sizes. Before committing, verify your GPU memory against the stated requirements: 80GB for gpt-oss-120b and 16GB for gpt-oss-20b, both under MXFP4 quantization.

## FAQ

### What is gpt-oss?

It is a series of open-weight language models from OpenAI, released as gpt-oss-120b and gpt-oss-20b, designed for reasoning, agentic tasks and developer use cases. The openai/gpt-oss repository holds reference inference implementations, while the weights are downloaded from Hugging Face.

### Are gpt-oss models free to use?

The models are released under the Apache 2.0 license, which the README describes as permissive with no copyleft restrictions or patent risk. The repository also contains a USAGE_POLICY file that the README does not explain, so review both before commercial deployment.

### Can I run gpt-oss locally?

Yes. The README states gpt-oss-20b runs within 16GB of memory and points to Ollama for consumer hardware, while gpt-oss-120b fits a single 80GB GPU such as an NVIDIA H100 or AMD MI300X. Both figures assume the MXFP4 quantization the models were post-trained with.

### How do I install gpt-oss?

The package requires Python 3.12 or newer. For serving, the README recommends installing a gpt-oss-specific vLLM build via uv from the wheels.vllm.ai index, then running vllm serve openai/gpt-oss-20b. The weights themselves come from Hugging Face.

### How do I use gpt-oss-20b with Ollama?

The README states that if you are trying to run gpt-oss on consumer hardware, you can use Ollama, and gives a command for it. It does not document the full Ollama workflow beyond that pointer, so check the linked guide for the current steps.

### How do I use gpt-oss-120b?

The README's Transformers example uses the model id openai/gpt-oss-120b with the pipeline API, and notes that the chat template applies harmony automatically. The README also states gpt-oss-120b fits into a single 80GB GPU such as an NVIDIA H100 or AMD MI300X.

## Sources

- [License: Apache-2.0](https://github.com/openai/gpt-oss/blob/main/LICENSE)
- [openai/gpt-oss on GitHub](https://github.com/openai/gpt-oss)
- [Project website](https://openai.com/open-models)
- [README](https://github.com/openai/gpt-oss/blob/main/README.md)
- [Releases](https://github.com/openai/gpt-oss/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/openai-gpt-oss
