Inside CollegesChat university-information: a survey-built record of college life
收集全国各高校招生时不会写明,却会实实在在影响大学生活质量的要求与细节
At a glance
- What is it?
- A Chinese community project that collects the university admission and life-quality details official pages leave out, then publishes them as a searchable site with a three-license split and an unusually candid disclaimer.
- Who is it for?
- The most useful thing in this repository is not the data volume but the shape of the disclaimer. CollegesChat states plainly that accuracy is not guaranteed, that listings may include unaccredited schools, that its own processing scripts may misclassify domestic institutions, and that it accepts no liability for a bad admissions decision.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- GitHub does not report a main language for this repository.
Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A dataset built around a gap in official admissions pages
The project description captures the premise in a single line: it collects the requirements and details that schools will not write down during admissions but that genuinely shape the quality of university life. That framing matters, because official admissions pages are written for a purpose. They must promote, they must be publishable by legal and marketing teams, and they must avoid the kind of frank operational detail that prospective students actually need in order to make a good decision.
The response to that gap was not a scrape and not an editorial team. It started as a survey shared through a Telegram channel, and the README credits two specific channel posts as its inspiration. The survey collected responses from people who already attend or recently attended each institution, which means the data carries a particular kind of authority: it is first-hand reporting from people with direct knowledge, gathered without a marketing filter.
The scale of that response shows in the repository numbers. The project has around 5,090 stars and 824 forks, with 35 open issues at the time of review. A star count on a data repository is a rough proxy for how many people plan to consult it during an admissions cycle, and for a fork count it suggests many readers want a local copy or a private variant. The default branch is `v2`, which signals that the current site generation path is a second generation, not a continuation of the first.
The published destination is colleges.chat, and the project is not archived. The most recent push recorded is 2026-09-21, which is recent enough to treat the repository as active work rather than a frozen archive. Someone is still changing something.
How the collection and publication pipeline works
The workflow is split across four separate channels, and the README routes each kind of participant to a different one. Reading an existing entry happens on the site itself. Filling out the survey happens at a separate submission host, submit.colleges.chat, kept deliberately off the main domain. Contributing data, asking a question, or querying records that are not fully public happens through GitHub Discussions.
Discussion-based participation is joined by a Telegram group for anyone who wants to ask about a specific institution or take part in project discussion. That gives the project three distinct conversation surfaces with three distinct audiences, which is a sensible design when the participants include students, prospective applicants, and outside contributors who each arrive with different questions.
The site itself is statically generated. The repository tree is small and mostly data: a `datas` directory holding the collected answers, a `docs` directory, a `questionnaire` directory, plus standard repository files. There is no application server in the tree. The generator that turns questionnaire data into a browsable site lives in separate repositories, which the README credits explicitly.
The credits table names four repositories and four distinct roles, and the split is the interesting part. This repository held the Mkdocs site generator and frontend in the older arrangement, and the CollegesChat/colleges.chat repository held the generated frontend source. The newer arrangement moves that work to CollegesChat.github.io for the Hugo frontend template and the website-generator repository for the Hugo site generator. The README marks the old pair as legacy and the new pair as current.
That is a full rewrite of the publishing stack, and it is the kind of change that quietly costs a project weeks. Moving from Mkdocs to Hugo meant migrating the generator, the templates, and the path structure at the same time as the data continued to grow.
$ ls CollegesChat/university-information
.github/
.gitignore
.gitmodules
LICENSE
datas/
docs/
questionnaireThe disclaimer is the most honest document in the repository
The disclaimer section is worth reading closely, because it documents failure modes the project has actually encountered rather than generic liability language. It starts by stating that content comes from the network and from survey collection, that accuracy cannot be guaranteed, and that the site should be treated as a reference to be cross-checked against other sources.
Then it names a specific risk. Not every school listed is necessarily a recognized institution in the official register. The listing may contain foreign universities and what the text calls diploma mills. Readers are told to determine for themselves whether a school has the credentials to admit students. The project then states that it accepts no responsibility if bad data on the site leads someone to choose the wrong school.
The third paragraph is a confession about the tooling. Some submitters did not fill in school names in a standardized format, and the data processing scripts may contain defects. As a result, some domestic universities may not be categorized into the correct province and city directory. This is a specific, falsifiable claim about a class of bug, and it tells a reader exactly where to look when a search returns nothing.
Three kinds of problem get three dedicated GitHub issue templates. A `malicious_data.yml` template handles infringement requests and untrue information. A `copyright_infringement.yml` template handles reports that the project itself is infringing someone's rights. A `document_correction.yml` template handles typos, stale facts, and omissions in the parts that are not questionnaire responses. Separating these matters because they have different urgency and different audiences.
CC BY-NC-SA 4.0
BSD 2-Clause
AGPL-3.0-or-laterRepository layout and where the real content lives
The repository tree has seven top-level entries and two of them are directories that hold the substance. The `datas` directory, confirmed by the default branch name `v2` pointing at its tree path, holds the collected survey answers. The `questionnaire` directory holds the survey definition, and the README links it as a separate repository reference, which suggests the questionnaire itself is maintained independently of the answers.
The `docs` directory is the remaining content directory. Given that the correction template is described as covering the parts that are not survey responses, `docs` is the most likely home for that hand-maintained or scraped material. That distinction has consequences for reuse, because the license terms differ between survey content and other documentation, and because corrections to `docs` follow a different path than corrections to a response.
The `.gitmodules` entry is the more structurally interesting one. A git submodule file means this repository references other repositories by commit reference rather than copying their contents. That fits the architecture described in the credits table: the generator and the frontend template live elsewhere and are pulled in at a fixed revision. It also means a build depends on submodule state, so a fresh clone needs submodule initialization before generation will work.
The license field in the repository metadata reads NOASSERTION, which is the GitHub label for a repository whose licensing cannot be reduced to a single standard identifier. Here that is accurate rather than evasive. The README documents three separate licenses covering three separate layers, and no single SPDX expression can describe that arrangement.
Why three licenses coexist in one project
The license section maps content to terms precisely, which is unusually careful. The questionnaire itself and the collected answers under `datas` are both under Creative Commons Attribution-NonCommercial-ShareAlike 4.0. The old Mkdocs site generator and frontend are under BSD 2-Clause. The new Hugo site generator and frontend are under AGPL-3.0-or-later.
The AGPL choice for the current generator is the one with real consequences. AGPL is a copyleft license written specifically so that modifications offered over a network must release their source. For a site generator that most people never modify and most people just run, that is a strong requirement. Combined with the noncommercial and share-alike terms on the survey data, the arrangement produces a clear split: read and browse freely, reuse the data only noncommercially with attribution and share-alike, and treat the current generator as copyleft.
Keeping the questionnaire and its answers under a noncommercial license is a defensible choice for a volunteer project. It prevents a commercial data product from being built directly on other people's unpaid answers, and it keeps any derivative share-alike so improvements propagate back into the commons.
Credits, sponsors, and the cost of running a public dataset
The acknowledgements section does two things. It thanks survey respondents and contributors, then it thanks developers across four repositories, with a note that internal repositories have been omitted from the list. That omission is disclosed rather than hidden, which is a small sign of editorial care.
The sponsor table is more revealing because it puts numbers on running costs. One sponsor contributed 11.35 USD for domain registration. A second sponsor contributed 22 USD, recorded as 21.68 USD actually spent, covering a domain renewal. That same sponsor contributed a server, with the stated purpose of improving site access speed. A third sponsor contributed 166.66 CNY, recorded as 23.18 USD actually spent, for another domain renewal.
Read together, those figures describe a site whose entire infrastructure is a domain name and a single server, funded by small individual donations and paid annually. The gap between pledged and actually spent amounts appears twice, which suggests a careful record of real costs rather than round numbers. The implied conclusion is that the operating cost of this kind of static dataset site is low enough to sustain through small recurring sponsorships, provided someone keeps renewing the domain.
That also sets the project's real fragility. A dataset this heavily dependent on volunteers, a small server budget, and a scattered multi-repository toolchain has no obvious successor. Its continuity depends on a handful of identifiable people continuing to show up, and the sponsor table names them.
Editorial conclusion
The most useful thing in this repository is not the data volume but the shape of the disclaimer. CollegesChat states plainly that accuracy is not guaranteed, that listings may include unaccredited schools, that its own processing scripts may misclassify domestic institutions, and that it accepts no liability for a bad admissions decision. Most dataset projects bury caveats in a footer. This one puts them in the README as the second section, before the credits.
For anyone deciding how to use it, the practical guidance is specific rather than general. Treat a school's page as a lead worth checking, then confirm accreditation status, tuition, and program details against official sources before acting. Use the three issue templates when something is wrong, because each one routes a different class of problem to a different reader. And if you intend to build on the work, read the license section first, because the questionnaire data, the old Mkdocs stack, and the new Hugo stack are covered by three different terms.
Frequently asked questions
What does CollegesChat/university-information actually collect?
It collects university admission requirements and life-quality details that schools do not publish during the admissions process but that shape day-to-day university life. The data originates from a survey distributed through a Telegram channel plus information gathered from the network, and it is published as a browsable site at colleges.chat. The repository itself holds the collected answers under the `datas` directory, not the site generator or frontend.
How does the project gather and publish its data?
Four channels handle different tasks. Reading existing entries happens on the site. Filling out the survey happens at submit.colleges.chat. Contributing data, asking questions, and partial lookups happen through GitHub Discussions. A Telegram group covers institution-specific questions and project discussion. Publication is static: the questionnaire data is turned into a site by a Hugo generator and Hugo frontend template that live in separate repositories, referenced here as submodules.
Can I rely on the school information without checking other sources?
Not on its own, and the project says so directly. The disclaimer states that content comes from the network and from surveys, that accuracy cannot be guaranteed, and that the site is a reference to cross-check. It warns that listed schools may include foreign institutions and unaccredited ones, and that the project accepts no liability if bad data leads to a wrong choice. It also notes that some domestic schools may sit under the wrong province or city directory because of nonstandard name entry and possible defects in the processing scripts.
How do I report incorrect or infringing content?
Three issue templates separate the cases. Use the malicious_data template for untrue information or infringement in the data. Use the copyright_infringement template to report that the project itself is violating someone's rights. Use the document_correction template for typos, outdated facts, and omissions in material that is not a questionnaire response. Splitting them matters because each has a different reader and a different urgency.
Which license applies to which part of the project?
Three licenses cover three layers. The questionnaire and the collected answers in `datas` are CC BY-NC-SA 4.0, meaning reuse must be noncommercial, attributed, and share-alike. The old Mkdocs generator and frontend are BSD 2-Clause. The current Hugo generator and frontend are AGPL-3.0-or-later. GitHub reports NOASSERTION for the repository because no single standard identifier covers that arrangement.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/collegeschat-university-information)