Model or dataset
OpenCSGs/csghub-server avatar
OpenCSGs/csghub-server

csghub-server: the Go backend behind a self-hosted model and dataset hub

csghub-server is the backend server for CSGHub which helps user to manage datasets, modes, and also run Model Inference, Finetune and Application Spaces.

1,051 stars233 forksGoApache-2.0

At a glance

What is it?
CSGHub Server is the REST API layer that handles users, organizations, model and dataset assets, LFS downloads and activity tracking for the CSGHub platform. It is a Go service you deploy with docker-compose, and it expects you to bring Gitea, Postgres, MinIO and an API token of at least 128 characters.
Who is it for?
Adopt csghub-server if you want a self-hosted catalogue for models and datasets with Git LFS storage, auto-tagging and activity tracking, and you are willing to run Postgres, Gitea and an S3-compatible store alongside it. Do not adopt it if you only need an artifact registry for a handful of files, or if you need model format conversion, which the roadmap still lists as unchecked.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What csghub-server actually manages

CSGHub Server is one component of CSGHub, which the README describes as an open source large model assets management platform. The server is the backend: it exposes a REST API and owns the records for users, organizations, models and datasets. The README's feature list is concrete about scope. It covers creation and management of users and organizations, auto-tagging of model and dataset labels, search across users, organizations, models and data, online preview of dataset files such as .parquet, content moderation for text and images, download of individual files including LFS files, and tracking of activity data such as download and like volume.

The intended user is an operator standing up an internal or public hub rather than an individual developer. The README points to the OpenCSG website as the hosted experience and offers csghub-server as the way to run that functionality yourself. That framing matters: this is infrastructure, not a library you import. If you are looking for a client to talk to an existing hub, this repository is the wrong side of the wire.

How the pieces fit: Gitea, Postgres, MinIO and the API layer

The architecture is visible in the repository layout and the compose file. The Go module is opencsg.com/csghub-server, and the dependencies show what the server leans on: gin for HTTP routing, pgx for Postgres, minio-go for object storage, and the DuckDB Go binding for querying data files. The README's extensibility section states that the server supports different git servers such as Gitea and GitLab, that LFS storage can be local or any S3-compatible service, and that content moderation can be enabled on demand against a third-party service.

So the data flow is split. Git repositories and LFS objects live behind a git server and an object store; the server keeps metadata, users, organizations and activity records in Postgres. The docker-compose.yml confirms this by declaring a Postgres image with the environment variable POSTGRES_MULTIPLE_DATABASES set to "starhub_server,gitea,mirror", which is the clearest statement in the repository that one database instance is expected to hold the server's own schema plus Gitea and a mirror schema. The compose file also declares memmachine services backed by pgvector and Neo4j, which sit outside the core asset-management path.

The trade-off is operational weight. Nothing in the README suggests a single-binary mode that avoids the git server and the object store. You are adopting a small constellation of services, and the server is the coordinator.

Installing csghub-server with docker-compose

The README's Quick Start gives a four-step path and states a system requirement of 4 CPU cores and 8GB of memory, tested in an Ubuntu 22 environment. Docker is a prerequisite you install yourself. The first command exports the API token, which the README says must be at least 128 characters long and is sent as a Bearer token on HTTP requests to the server.

bash
export STARHUB_SERVER_API_TOKEN=<API token>
mkdir -m 777 gitea minio_data
curl -L https://raw.githubusercontent.com/OpenCSGs/csghub-server/main/docker-compose.yml -o docker-compose.yml
docker-compose -f docker-compose.yml up -d

The mkdir step is not incidental. It creates the gitea and minio_data directories with permissions 777 before the compose file mounts them, so the containers can write to host paths. If you skip it or run it from a different working directory, the mounts in docker-compose.yml will point at paths that do not exist.

For development against the source tree, the README shows starting a service directly with the Go tool and a TOML config file. The example config lives at common/config/config.toml.example, and the README says all available settings are defined in common/config/config.go, with snake_case keys mapping to struct fields.

bash
go run cmd/csghub-server/main.go start server --config local.toml
go run cmd/csghub-server/main.go deploy runner --config local.toml

The Makefile adds the pieces you need after a config exists. Database migrations run through the migration subcommand, and the Makefile wraps it as migrate_local and db_migrate against local.toml. A build target produces ./bin/csghub-server from ./cmd/csghub-server.

bash
make build
make migrate_local

After `make migrate_local`, the schema in the starhub_server database should be current. The README does not describe what the server logs on a successful boot, so treat a clean migration exit as your checkpoint rather than a specific log line.

Where csghub-server is the wrong tool

The README is candid in one place that matters: the roadmap lists model format convert as unchecked, while every other listed item is checked. If your workflow depends on converting between model formats inside the hub, that capability is not there yet according to the project's own roadmap, and you would be running a separate conversion step outside it.

The second limitation is dependency surface. The compose file pulls images from multiple registries, including a Postgres build from a regional Aliyun registry and memmachine images from Docker Hub. In a network-restricted environment, that mix is a real procurement problem, and the README does not document an offline or air-gapped installation path. It also does not document rollback beyond the db_rollback Make target, which invokes the migration rollback subcommand; there is no described procedure for reverting a container image upgrade.

Third, this is not a lightweight file server. If you want to publish three models to a handful of colleagues, the combination of Postgres, a git server, an object store and a moderation service is more moving parts than the problem deserves. A plain S3 bucket plus a README would serve that case with less to operate.

How it differs from Git LFS on a plain git host

The obvious alternative is running Git LFS directly on a git host such as Gitea or GitLab and skipping csghub-server entirely. The difference in approach is where the metadata lives. With plain Git LFS, the repository is the unit of record: a model is a directory in a repo, and its description is whatever you write in a markdown file. Search, tagging and activity counts are whatever your git host happens to provide.

csghub-server puts a database-backed catalogue in front of that. The README's feature list is essentially the delta: auto-tagging of model and dataset labels, search across users, organizations, models and data, online preview of .parquet files, content moderation, and tracking of downloads and likes. Those are catalogue concerns, and they need a schema, which is why Postgres is in the compose file alongside the git server rather than replaced by it. The README states the server supports Gitea and GitLab as backends, so the git host is not eliminated either way; it is demoted to a storage detail.

If your team already runs GitLab with LFS and nobody needs faceted search or download statistics, csghub-server adds a database and an API surface for capabilities you are not using. If people keep asking which model is the current one and how many times it has been pulled, the catalogue is the point.

Maintenance, releases and licence

The repository is not archived, and the last push was on 2026-08-09, which is the same timestamp as the v2.4.0-ce release. Before that, v2.3.0-ce landed on 2026-07-06 and v2.2.0-ce on 2026-06-09, so the release cadence across those three versions is roughly monthly. The -ce suffix on every release tag suggests a community edition line, though the README does not explain what an enterprise edition would add.

Upgrade cost is mostly in the migrations. The Makefile provides db_migrate and db_rollback against local.toml, which means schema changes are handled by the server's own migration command rather than by hand. What the README does not describe is whether migrations are additive or whether a rollback after a release upgrade leaves data intact; the db_rollback target exists, but its semantics are not documented in the README.

Licensing is Apache-2.0, stated in the README and present as the LICENSE file at the repository root. That is a permissive licence with an explicit patent grant, and it does not impose copyleft on your own code. It does not, by itself, settle the licences of the container images the compose file pulls, which come from separate projects and registries. Check those separately before redistributing a bundled deployment.

Editorial conclusion

Adopt csghub-server if you want a self-hosted catalogue for models and datasets with Git LFS storage, auto-tagging and activity tracking, and you are willing to run Postgres, Gitea and an S3-compatible store alongside it. Do not adopt it if you only need an artifact registry for a handful of files, or if you need model format conversion, which the roadmap still lists as unchecked. Before committing, verify that your team can supply the STARHUB_SERVER_API_TOKEN at the documented length, that the ports in docker-compose.yml do not collide with services you already run, and that the Apache-2.0 LICENSE file matches your redistribution plans.

Frequently asked questions

What is csghub-server and what does it do?

It is the backend server for CSGHub, an open source platform for managing large model assets. According to the README, it handles users and organizations, models and datasets, auto-tagging, search, dataset preview, content moderation, LFS file downloads and activity tracking through a REST API.

How do I install csghub-server?

The README's Quick Start deploys it with docker-compose after exporting STARHUB_SERVER_API_TOKEN, creating the gitea and minio_data directories, downloading docker-compose.yml and running docker-compose up -d. It states a requirement of 4 CPU cores and 8GB of memory and says the project was tested in Ubuntu 22.

What are the requirements for the csghub-server API token?

The README says the API token should be at least 128 characters long and that HTTP requests to csghub-server must send it as a Bearer token for authentication. It is supplied through the STARHUB_SERVER_API_TOKEN environment variable.

Does csghub-server support model format conversion?

Not according to its own roadmap. The README lists model format convert as the only unchecked item on the roadmap, while items such as Git LFS, dataset preview, auto-tagging, S3 support and one-click model deploy are marked complete.

Which git servers and storage backends does csghub-server support?

The README states that it supports different git servers such as Gitea and GitLab, and that LFS storage can be local or any third-party cloud storage service compatible with the S3 protocol. The docker-compose.yml ships with a MinIO data directory alongside the Postgres and Gitea services.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/opencsgs-csghub-server.svg)](https://hysenlabs.com/projects/opencsgs-csghub-server)