Open-source project
ckan/ckan avatar
ckan/ckan

CKAN: the AGPL data portal platform behind national open data sites

CKAN is an open-source DMS (data management system) for powering data hubs and data portals. CKAN makes it easy to publish, share and use data. It powers catalog.data.gov, open.canada.ca/data, data.humdata.org among many other sites.

5,127 stars2,116 forksPythonNOASSERTION

At a glance

What is it?
CKAN is a Python data management system for publishing and cataloging datasets, with a full API and a plugin extension model. It is the right tool for institutional portals, and the wrong one for a weekend catalog.
Who is it for?
Adopt CKAN if you are standing up an institutional or government data portal that needs a catalog, a data API and a plugin surface, and you have the Python and PostgreSQL skills to run it. Do not adopt it for a small static catalog of a few files, where a generated site is cheaper to operate.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What CKAN solves, and who is actually adopting it

A data portal is not a file server. The moment an organization publishes more than a handful of datasets, someone has to answer questions the filesystem cannot: which department owns this table, when was it last refreshed, what licence covers it, which of the twenty CSVs in this release is the current one, and how does an external application fetch it without a human clicking through pages. CKAN exists to answer those questions as structured metadata rather than as prose on a page.

The project describes itself as a data management system that provides a platform for cataloging, storing and accessing datasets, with a rich front-end, a full API for both data and catalog, and visualization tools. The README names catalog.data.gov, open.canada.ca/data and data.humdata.org as sites it powers. That list is the clearest statement of the target user: national statistics offices, humanitarian coordination bodies, municipal open data teams, research institutions. These are organizations with a publishing mandate, multiple contributing departments, and a need for machine-readable access.

The topics on the repository include digitalpublicgoods and dpg, alongside sdg16. That framing matters for procurement. CKAN is positioned as public infrastructure, not as a commercial product with a support contract attached, and the support channels in the README reflect that: a ckan-dev mailing list, a Gitter chat, and the CKAN tag on Stack Overflow. If you need a vendor to call at 2 a.m., the README does not point you at one.

Who it is not for: a solo developer who wants to list five PDFs. The stack below is real infrastructure with real operational cost, and the return only appears once you have enough datasets and enough consumers to justify it.

The architecture: Flask, PostgreSQL, Solr and a plugin registry

The dependency list tells you most of the shape. requirements.txt pins flask 3.1.3, flask-login, flask-session, flask-wtf and jinja2 for the web layer, sqlalchemy with alembic for persistence and migrations, and rq with croniter for background jobs. The presence of rq and croniter is the giveaway that some work in CKAN does not happen inside the request cycle: harvesting remote catalogs, indexing, and scheduled tasks run as jobs against a queue.

Search is not served by the database. The repository layout includes a ckanext directory with a multilingual/solr subtree, and setup.py declares message extractors for ckanext/multilingual/solr/*.txt. Solr is therefore a separate service in the deployment, not a library you import, which means a CKAN install has at least three moving parts to keep alive: the web application, PostgreSQL and Solr.

The extension model is the part that determines whether CKAN fits your organization. ckanext is a first-class directory in the repository, and pyproject.toml registers pytest markers named provide_plugin and with_plugins, which the test suite uses to enable plugins during tests. Extensions are Python packages that hook into the core, so a custom metadata schema, a new harvester or an authentication backend is an extension rather than a fork. That is a genuine architectural commitment, and it is why the ckanext namespace has accumulated so much third-party code.

The front-end is not a JavaScript SPA. package.json lists bootstrap 5.3.6, jquery 3.7.1, htmx.org 2.0.10, select2 and moment, with gulp for asset building and sass for stylesheets. Server-rendered templates with progressive enhancement, built by gulp rather than a modern bundler. Node is pinned to >=24 <25 in the engines field, so the asset toolchain has its own version constraint separate from the Python side.

Installing CKAN and publishing a first dataset

The README does not contain installation steps. It says plainly: see the CKAN Documentation for installation instructions. That is the authoritative source, and it is where you should start, because a CKAN deployment involves PostgreSQL, Solr, a Python virtual environment and a WSGI server, and the documentation covers the supported combinations.

The repository does expose the pieces you will assemble. The WSGI entry point is wsgi.py at the top level, and ckan-uwsgi.ini is the sample uWSGI configuration shipped with the source. A production deployment points a WSGI server at the former using the latter as a starting point.

On the Python side, setup.py declares two extras groups, requirements and dev, read from requirements.txt and dev-requirements.txt. That means the package metadata already distinguishes a runtime install from a development install, and you should not need to guess which files to install from.

The front-end assets are not built by pip. package.json defines the gulp tasks, and the postinstall hook runs gulp updateVendorLibs automatically, so a plain npm install pulls the vendor libraries into place. The build task compiles the stylesheets and scripts:

bash
npm install
npm run build

After that, the configuration file is the next thing to understand. The test suite is configured through a --ckan-ini flag pointing at test-core.ini, declared in the pytest addopts in pyproject.toml. That is the pattern for the whole application: an ini file supplies the database URL, the Solr URL and the site settings, and every entry point takes a path to one.

bash
pytest --ckan-ini=test-core.ini

Running that command is how you confirm the install before touching a browser. The markers registered in pyproject.toml (ckan_config, provide_plugin, with_plugins) are what the fixtures use to patch configuration and enable plugins per test, so a passing run means the database, the configuration loader and the plugin registry are all wired correctly. Only then does it make sense to create an organization and a dataset through the web interface or the API.

Where CKAN becomes the wrong tool

The licence is the first constraint, and it is not a small one. CKAN is released under the GNU Affero General Public License v3.0, stated in the README and in the LICENSE.txt file at the repository root. The AGPL's network clause is the relevant part for a portal: if you modify CKAN and let users interact with it over a network, the licence's terms reach that modified version. Organizations that treat their portal code as proprietary should read the licence text and talk to counsel before building on it. This is not a legal opinion, and the README does not attempt to soften the point.

The second constraint is operational surface. PostgreSQL, Solr and a Python application with a background job queue is a stack that needs someone who can upgrade three things in step. The repository shows three maintained release lines at once (2.12.0, 2.11.6 and 2.10.11, all dated 2026-08-26), which is good for stability but also means you must decide which line you are on and how long you will stay there.

The third is the front-end toolchain. Node is pinned to >=24 <25. If your build environment ships an older Node, the asset pipeline will not run, and because assets are built by gulp rather than installed from a registry, you cannot skip that step. A team that has standardized on an older Node LTS will feel this immediately.

Finally, consider the failure mode of the extension model. Extensions hook into core, so a major version upgrade can break third-party ckanext packages you depend on. The repository's changes directory and towncrier configuration in pyproject.toml define a migration notes category precisely because upgrades sometimes require action. If your portal depends on an extension that is not tracking the current release line, your upgrade path is blocked by someone else's maintenance schedule, not your own.

CKAN compared with a static catalog generator

The realistic alternative for a small publisher is not another portal platform. It is a static site generator plus a JSON or CSV index: you commit dataset metadata to a repository, a build step renders HTML, and the files sit in object storage behind a CDN. There is no PostgreSQL, no Solr, no job queue and no server to patch.

The difference in approach is where the metadata lives and who can change it. In a static catalog, metadata is a file in version control, and publishing means a pull request. In CKAN, metadata is rows in PostgreSQL managed through a web interface and an API, and publishing is a workflow that non-developers can complete without touching git. That is the trade: CKAN buys you delegated publishing, search across a large corpus, and a data API that external applications can query, at the cost of running a stateful service.

There is a second difference worth naming. A static catalog gives you no harvest endpoint and no standard catalog API, so other portals cannot pull your metadata automatically. CKAN's API and its harvester extensions are what make it a node in a network of catalogs rather than an isolated website. If federation with other portals is a requirement, the static approach does not have an answer.

A third option, a general-purpose content management system with a file field, is the one to be most suspicious of. It will store your datasets, but it will not give you dataset-level metadata, a catalog API, or a search index that understands the difference between a dataset and a resource inside it. You will end up reimplementing CKAN's core model badly.

Maintenance, releases and upgrade cost

The repository is not archived, and the last push was on 2026-09-22. Three release lines were published on the same day in August 2026: ckan-2.12.0, ckan-2.11.6 and ckan-2.10.11. Maintaining three lines simultaneously is a deliberate policy, and it tells you the project expects deployments to lag. You are not required to be on the newest version the week it lands.

Upgrade cost is where the extension model bites. The towncrier configuration in pyproject.toml defines five change categories: migration notes, major features, minor changes, bugfixes, and removals and deprecations. The existence of a dedicated migration notes category, and of a removals and deprecations category, means the changelog is structured to tell you when an upgrade requires action rather than merely adding features. Read the migration notes for your target version before upgrading, and check the deprecations category for anything your extensions rely on.

On the Python side, requirements.txt is autogenerated by pip-compile from requirements.in, with Python 3.10 noted in the file header. That means the pinned versions are a tested set, not a loose range, and upgrading dependencies should go through requirements.in and a recompile rather than editing the lock file. The pyproject.toml pyright configuration targets pythonVersion 3.10 as well, so type checking and dependency resolution agree on the interpreter.

Licence implications for a deployment: the AGPL v3.0 covers CKAN itself, and the README states the copyright runs from 2006 to 2023 for the Open Knowledge Foundation and contributors. Extensions you write are your own code and their licensing is your decision, but the boundary is worth clarifying with counsel before you ship. The README directs security reports to [email protected] rather than public GitHub issues, which is the channel to use if you find a vulnerability.

Editorial conclusion

Adopt CKAN if you are standing up an institutional or government data portal that needs a catalog, a data API and a plugin surface, and you have the Python and PostgreSQL skills to run it. Do not adopt it for a small static catalog of a few files, where a generated site is cheaper to operate. Before committing, verify the installation path in the CKAN documentation for your target version, confirm the AGPL v3.0 obligations with your legal team, and check which of the three maintained release lines (2.12.0, 2.11.6, 2.10.11) your deployment will track.

Frequently asked questions

What does CKAN mean?

The README does not expand the acronym anywhere; the project presents itself only as CKAN, an open source data management system for powering data hubs and data portals. Its formal description is 'CKAN: The Open Source Data Portal Software'.

Does CKAN still work?

The repository is not archived and the last push was on 2026-09-22, with three release lines published on 2026-08-26: ckan-2.12.0, ckan-2.11.6 and ckan-2.10.11. The README also lists catalog.data.gov, open.canada.ca/data and data.humdata.org as sites it powers.

How safe is CKAN?

The README asks that potential security vulnerabilities be emailed to [email protected] rather than filed as public GitHub issues, and the repository carries a SECURITY.md file at the top level. No audit report or vulnerability history is published there, so no broader safety claim can be drawn from it.

Official sources

  1. ckan/ckan on GitHub
  2. Issues
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ckan-ckan.svg)](https://hysenlabs.com/projects/ckan-ckan)