Self-hosted service
digoal/blog avatar
digoal/blog

digoal/blog: a Chinese PostgreSQL writing archive kept in Git

AI,Opensource,Database,Business,Finance,Minds. git clone --depth 1 https://github.com/digoal/blog

8,577 stars1,923 forksHTMLGPL-2.0

At a glance

What is it?
Not a program but a document tree. Thousands of Chinese articles on PostgreSQL, Greenplum and, increasingly, AI search, filed into month folders and 36 topic indexes under a GPL-2.0 license.
Who is it for?
The honest case for this repository is narrow and strong. If you work with PostgreSQL or Greenplum and you read Chinese, it is one of the larger openly licensed collections of production database writing available anywhere, and cloning it is the only practical way to search it.
Can I use it commercially?
Yes, with conditions. GPL-2.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 20 days ago.
What is it written in?
Mainly HTML, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

This is a document archive rather than a package

The first thing to settle is what kind of thing this is. There is no build, no package manifest, no test suite and no executable anywhere in it. GitHub reports the dominant language as HTML, which comes from the two group chat screenshots the README embeds, not from code. The repository description is a run of subject tags, AI, Opensource, Database, Business, Finance, Minds, followed by a shallow clone command. That is how someone describes a pile of writing they intend to publish, not software they intend to ship.

So there is nothing to install. What there is, instead, is a very large body of Chinese technical prose, and the whole point of putting it in Git is that it sits in one place under one license rather than spread across a personal site that can rot. The homepage field points at the repository's own README on GitHub, which tells you there is no separate site to visit. The clone command exists so people can mirror the archive, not so they can run it.

The topic tags make the focus legible without reading any of the body: database, postgres, postgresql, pg, pgsql, greenplum, gpdb, hawq, enterprisedb, oracle, mysql and mongodb. Eleven of those twelve are database engines or vendor names, which is an accurate summary of where the writing goes.

The README is three reading lists before it is an index

The README opens with a one line link to an about page at me/readme.md, then runs a numbered list of thirty three video and course series. The earliest entries are classroom recordings, a four day and a five day course on PostgreSQL 9.3 administration and optimisation, a three day optimisation course, a one day course on 9.1, and a lecture series. The list then runs forward through 2017, 2018, 2019, 2020, 2021 and into 2026, moving from migration training to architecture material to source code reading.

The video downloads all point at one Baidu pan link with an extraction code, and the README asks readers to report it if the link stops working. That single point of failure is worth knowing about before you plan any course from this list.

A second list of nineteen items follows under a study materials heading. These are reference style rather than video: a database security guide, a set of PostgreSQL conventions, a frequently used SQL collection, an eight part quick start on application development and administration that runs from building a local environment through transactions and locking to stored procedures and triggers, and two selection guides written for people choosing a database. Two further sections cover gratitude and a thinking column of seventeen pieces, where the fourth section of the README links straight to class/35.md, the same file that appears as category 35 in the index further down.

Thirty six category files are the only working index

After a contact table pointing at a WeChat account and a DingTalk group, the README presents the navigation layer: thirty six bold links, class/1.md through class/36.md, laid out four to a row. The categories are application development, day to day maintenance, monitoring, backup restore and disaster recovery, high availability, security and auditing, problem diagnosis and performance optimisation, streaming replication, read write separation, horizontal sharding, OLAP and MPP, database extensions, new features by version, kernel internals and development, classic case studies, HTAP, stream computing, time series and spatial data, graph style search, GIS, Oracle compatibility, database selection, benchmarking, working practices, DaaS, vertical industry applications, standardisation, version upgrades, homogeneous and heterogeneous data synchronisation, data analysis, course series, miscellaneous, recruitment and job seeking, conferences and training, thinking pieces, and video recordings.

That taxonomy is the most useful thing in the repository and it is worth reading twice, because it encodes how a working database team actually splits its reading. Some of the categories describe a subsystem, some describe a lifecycle stage, and some describe a job function.

The README is also candid about the coverage gap. A note above the table says that most articles are not categorised and directs readers to the final section instead. There is no search box, no tag file and no front matter in the documents themselves, so the chronological list at the bottom is a wall of links and the category files are the only structured view. Anyone planning to mine this archive should clone it and grep rather than browse.

Month folders, and a date scheme that does not always hold

Documents follow a predictable path and filename. The chronological list at the end of the README runs newest first, with entries such as 202609/20260916_01.md and 202609/20260915_01.md, so the folder is a year and month and the filename repeats that prefix with a day, an underscore and a two digit sequence number. Articles published on the same day get 01, 02, 03 and onward. The tree begins at .gitignore and then runs through month folders, starting at 197001 and 197002 and continuing by month through the rest of the archive.

That 197001 folder is a trap for anyone writing a script over the archive. It is not a publication date. It holds entries such as 20190214_01.md and 20200804_01.md, which means undated or awkward material was parked in a month that cannot exist and never moved. If you sort by folder you will push several years of writing to 1970, and if you build a timeline from directory names it will be wrong.

The date inside the filename is the one to trust. Parsing the filename rather than the directory gives a correct chronological ordering, with the parked articles falling back to whatever month they were eventually filed under.

PostgreSQL is the constant and AI search is the new growth

The subject matter has visibly widened. The newest listed entry, 202609/20260916_01.md, closes the eighth episode of a live stream marking thirty years of PostgreSQL, which tells you the archive still treats the database itself as its anchor. The thinking column beside it runs to at least 534 instalments under a recurring series name, and covers how to read source code with AI assistance rather than only what the source code does.

The clearest signal of the shift sits in late 2025, where a run of source code reading series appeared back to back: OceanBase across twenty four instalments, LangChain across twenty one, DuckDB across twenty seven, Milvus across thirty one, VectorChord across thirty nine, pgvector across twenty seven, DuckPGQ across thirty five, the streaming disk based implementation behind pgvectorscale across thirty nine, pg_tokenizer across twenty nine, a bm25 implementation for VectorChord across twenty, and SeekDB across thirteen. A university course on large model hands-on work with Dify and AI search follows, covering vector, keyword, scalar, graph and hybrid retrieval in one syllabus. Near the top of the current list sit an interpretation of a survey on agent memory and a piece reporting that sixty researchers including groups at Stanford and Google argue against treating vector retrieval as long term memory.

Read together, those entries describe a shift from a database specific archive to a database plus AI archive, with PostgreSQL as the connecting thread and vector search as the current interest. It is the single most useful thing to know before deciding whether this repository is worth your time.

GPL-2.0, three tags, and what this repository cannot do

The license is GPL-2.0, and applied to prose rather than code it has a specific shape. You may copy and redistribute the articles, including commercially, provided the license travels with the work and the source is credited. The README asks for attribution on repost, which lines up with the license terms rather than contradicting them. Because there is no code here, the copyleft has nothing to propagate into a linked work, which makes the practical obligation simple: keep the license, credit the author.

The release list is thin and not what releases usually mean. There are three published tags, 20250207, 20240202 and 20230824, all named after dates, all with empty notes. The most recent tag is dated 2025-02-07 while the last push was on 2026-09-17, so tagging stopped being a habit some time ago and there is no versioned changelog to read. The repository is not archived.

The limits are worth stating plainly. Everything is in Chinese, with no translation directory and no build step that produces one. There is no search interface, no feed file in the tree, and no per article metadata beyond the filename. The 130 open issues are a support surface for a repository that mostly holds text, and the chronological list in the README is long enough that browsing it is not a realistic discovery method. For a reader who does not read Chinese, the practical value of the whole archive is close to zero regardless of how good the writing is.

Editorial conclusion

The honest case for this repository is narrow and strong. If you work with PostgreSQL or Greenplum and you read Chinese, it is one of the larger openly licensed collections of production database writing available anywhere, and cloning it is the only practical way to search it. What it does not give you is a database, an English edition, a working index or a version history worth the name. The 36 files under class/ are the whole navigation layer, the most recent entries in the chronological list are AI search and agent memory material rather than database internals, and the three published tags are date stamps with no notes attached. Start at class/7 for troubleshooting and performance, class/12 for extensions, class/24 for working practices, or clone the archive and grep it for the specific error message you are stuck on.

Frequently asked questions

What is the digoal blog GitHub repository?

It is a GPL-2.0 archive of Chinese language technical articles, mostly about PostgreSQL and Greenplum, published as Markdown in month named folders using a YYYYMM/DD_NN.md path. It holds no software. GitHub reports HTML as the dominant language because of two embedded screenshots, and the homepage field points back at the README inside the repository, so the GitHub copy is the publication.

Can I install or run anything from the blog repository?

No. There is no package manifest, no build script and no executable in it. The clone command in the repository description exists so that people can mirror the archive. What `git clone --depth 1 https://github.com/digoal/blog` gives you is a directory of Markdown files to read or search, and the 36 files under class/ are the closest thing the project has to an index.

Is there an English version of the blog, and how do I find a specific topic?

There is no English edition and no translation layer: no second directory, no build step, nothing generated. Finding a topic means picking a category file such as class/7.md for diagnosis and performance or class/12.md for extensions, or cloning the repository and searching the text yourself. The chronological list at the bottom of the README is a wall of links rather than a usable index.

What license is the blog content under and can I republish it?

The repository is licensed GPL-2.0. For prose that means you may copy and redistribute the articles, including commercially, as long as you keep the license and credit the source, which is what the README asks for as well. The three published tags, 20250207, 20240202 and 20230824, are date stamps with no release notes attached rather than versions of anything.

Official sources

  1. digoal/blog on GitHub
  2. License: GPL-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/digoal-blog.svg)](https://hysenlabs.com/projects/digoal-blog)