replicate/cog: packaging machine learning models as Docker images
Containers for machine learning
At a glance
- What is it?
- Cog turns a cog.yaml file and a Python runner class into a production Docker image with CUDA, Python and dependencies resolved, plus a generated HTTP prediction endpoint. It fits researchers shipping a model, not teams that need a hand-written Dockerfile.
- Who is it for?
- Cog suits a researcher or small team that has a model working in a notebook and needs it behind an HTTP endpoint without writing a Dockerfile, especially if the target is Replicate or any host that runs Docker images. It is the wrong tool if you need a hand-tuned image, a non-Python serving stack, or native Windows support, since the README states Cog does not support Windows natively and points to WSL 2.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 15 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Cog solves for people shipping a model
The README states the problem plainly: it is hard for researchers to ship machine learning models to production, and Docker is part of the solution but painful to get right. Dockerfiles, pre- and post-processing, Flask servers and CUDA versions all have to line up, and the README says the researcher often ends up sitting down with an engineer to get the model deployed. Cog targets that gap. It is for someone who has a working model in Python and wants a container that runs it, without becoming the person who maintains a Dockerfile. The audience is narrow on purpose: model authors, not platform teams. The README says deployment can go to your own infrastructure or to Replicate, and the project was created by Andreas Jansson and Ben Firshman, with the README noting Firshman created Docker Compose and Jansson built deployment tooling at Spotify. That history explains the design bias. Cog assumes you already know how your model runs and only want the environment and the serving layer handled for you.
cog.yaml, the runner class and the generated OpenAPI schema
Cog splits a model into two declarations. cog.yaml describes the build environment: whether a GPU is needed, system packages, the Python version, and a requirements file. The README example sets gpu: true, installs libgl1 and libglib2.0-0, pins python_version to 3.13, and points python_requirements at requirements.txt. The second declaration is the runner itself, referenced by a run key such as run.py:Runner. The runner subclasses BaseRunner, loads weights in a setup method, and exposes a run method whose arguments are typed with Input and annotated with Python type hints such as Path. Those hints are the interface contract. The README says Cog uses the model's types to generate an OpenAPI schema, validate inputs and outputs, and dynamically generate a RESTful HTTP API served by a Rust/Axum server. So the data flow runs one way: Python type annotations become a schema, the schema becomes validation, and a compiled server exposes it over HTTP. That is the whole trick, and it is why the README claims no more CUDA hell. Cog holds a compatibility table for CUDA, cuDNN, PyTorch, TensorFlow and Python combinations and picks one for you. The cost is that you are trusting that table instead of controlling the base image.
Installing Cog and running a first prediction
Cog is a Go binary, distributed through a Homebrew tap, an install script at cog.run, and direct release downloads. On macOS the README gives brew install replicate/tap/cog as the easiest route. The install script works across shells, and there is a manual path that curls the release asset matching your OS and architecture into /usr/local/bin. Docker is a prerequisite on every platform, and the README adds that if you install Docker Engine rather than Docker Desktop you also need Buildx. Windows is not supported natively; the README directs Windows 11 users to WSL 2 and then the Linux instructions.
brew install replicate/tap/cogOnce the binary is on PATH, the workflow is the one from the README. You write cog.yaml, write the runner, then run the model locally against a real input. The -i flag passes an input value, and the @ prefix reads a file from disk, so [email protected] sends the local file as the image argument.
cog run -i [email protected]The README's sample output shows the build step, the run step, and a line reporting that output was written to output.jpg. For a served endpoint instead of a one-shot run, cog serve builds the environment and starts the HTTP server on a port you choose, which is useful for testing the generated API before deploying.
cog serve -p 8080With the server up, predictions go to the /predictions path as JSON, with the model inputs nested under an input key. The README shows this curl call against port 5000, the port used by the built image's default server.
curl http://localhost:8080/predictions -X POST \
-H 'Content-Type: application/json' \
-d '{"input": {"image": "https://.../input.jpg"}}'For deployment, cog build -t my-classification-model produces a normal image, and the README runs it with docker run -d -p 5000:5000 --gpus all. Nothing about that container is Cog-specific at runtime, which is the point.
Where Cog gets in the way
The generated environment is the product, and it is also the constraint. If your model needs a base image Cog does not produce, a system library installed in an unusual way, or a serving framework other than the generated Axum server, you are working against the tool rather than with it. The README's own framing is a trade: you give up control of the Dockerfile and get a curated one. Teams with an existing image pipeline and a reason to keep it will find Cog redundant. There is a second, quieter limit. The Python package metadata in pyproject.toml lists requires-python as >=3.10 with classifiers through 3.13, so the supported interpreter range is bounded, and cog.yaml pins a single python_version per project. A model that needs two interpreters, or a version outside that range, has no path through the config as shown. Platform support is the third edge: the README states Cog does not natively support Windows, so a Windows-only team is running WSL 2 whether or not that fits their deployment story. Finally, the README notes that development environment instructions for cog exec and the Jupyter notebook workflow were intentionally left out of the README for now. Those commands exist in the repository, but a reader relying on the README alone will not find them, which is a documentation gap worth knowing about before you plan around them.
Cog against a hand-written Dockerfile and BentoML
The honest alternative for many teams is the thing Cog is trying to replace: a Dockerfile you write yourself, plus a small web framework for the endpoint. That approach gives you total control of the base image, the layer order, the entrypoint and the server, and it costs you the compatibility work Cog does for you. The difference is not capability but where the knowledge lives. With a Dockerfile, the CUDA and PyTorch pairing is your problem and your documentation. With Cog, it is a table the project maintains and you inherit. BentoML is the other comparison worth drawing, since it also targets model serving in Python. BentoML's model is a bento: a build artifact that bundles a service definition, dependencies and model files, with its own runner abstraction and its own deployment targets. Cog's model is narrower and more literal. It produces a Docker image, uses Python type hints to derive an OpenAPI schema, and serves through a Rust/Axum server rather than a Python web stack. If your team already runs BentoML, moving to Cog means giving up that ecosystem for a smaller, container-first abstraction. If you have never needed a serving framework and only want a container, Cog's narrower scope is the reason to pick it.
Maintenance, releases and the Apache-2.0 licence
The repository is not archived, and the last push was on 2026-09-10, which is recent. Recent releases are v0.22.0 on 2026-08-14, v0.21.0 on 2026-06-16 and v0.21.0-rc.3 on 2026-06-05. The version numbering is pre-1.0 and the release cadence shows release candidates appearing before stable tags, so pinning a specific version rather than tracking latest is the safer default. Upgrade cost is mostly the image rebuild. Because cog.yaml pins a Python version and a requirements file, a new Cog release can change the base image and the resolved dependency set underneath you, which means a version bump is a rebuild-and-retest event, not a no-op. The Python package declares coglet>=0.1.0,<1.0 as a dependency, so the SDK side has its own version floor to watch. On licensing: the repository is Apache-2.0, which permits commercial use and modification with the usual notice and patent terms. The Python package metadata points its license at the LICENSE file rather than declaring an SPDX string in pyproject.toml. That is a packaging detail, not a legal problem, but if your organisation scans licences automatically, confirm what your tooling reads from the wheel. Nothing here is legal advice; read the LICENSE file before you rely on it.
Editorial conclusion
Cog suits a researcher or small team that has a model working in a notebook and needs it behind an HTTP endpoint without writing a Dockerfile, especially if the target is Replicate or any host that runs Docker images. It is the wrong tool if you need a hand-tuned image, a non-Python serving stack, or native Windows support, since the README states Cog does not support Windows natively and points to WSL 2. Before adopting it, check the cog.yaml keys your model needs against the docs, confirm your CUDA and PyTorch combination is one Cog knows, and run cog run -i on a single example input to see what the generated OpenAPI schema does with your types.
Frequently asked questions
What is Cog in Python?
Cog is an open-source tool from Replicate that packages machine learning models in a standard, production-ready container. Its Python side is the runner class you write, subclassing BaseRunner with a setup method and a run method whose typed arguments become an OpenAPI schema. The cog package on PyPI is that SDK, and the cog command itself is a Go binary.
Is Replicate free to use?
The README does not describe Replicate's pricing. It says you can deploy a packaged model to your own infrastructure or to Replicate, and that is the extent of what the repository material covers.
What are some good alternatives to Replicate?
The README does not list alternatives to Replicate as a hosting service. On the tooling side, the closest comparison the repository supports is a hand-written Dockerfile plus your own web framework, since Cog's stated goal is to replace that work with a generated image and server.
How do I get an API key for Replicate?
The README does not cover API keys. It documents running a model locally with cog run, cog serve and the /predictions endpoint, and mentions Replicate only as a deployment target.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/replicate-cog)