NLP-Knowledge-Graph: A Chinese Research Collection on NLP, KG, and LLMs
自然语言处理、知识图谱、对话系统,大模型等技术研究与应用。
At a glance
- What is it?
- lihanghang/NLP-Knowledge-Graph is a curated research notes repository created in August 2019 that gathers paper analyses, conference listings, tool comparisons, and dataset references for Chinese NLP, knowledge graph construction, QA systems, and large language model research. It is a reading and reference resource, not a software library.
- Who is it for?
- NLP-Knowledge-Graph is useful for Chinese-speaking researchers and engineers who want a single starting point for knowledge graph and NLP literature, curated with summaries and mind maps rather than raw paper lists. It is not a software package, and it does not cover recent LLM benchmarks or transformer architectures released after its last active update cycle.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 28 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A Research Notes Repository, Not a Software Library
lihanghang/NLP-Knowledge-Graph is a collection of reading notes, paper summaries, and external resource links for researchers working on natural language processing, knowledge graphs, and dialogue systems. The README, created on 2019-08-24, positions it as a series of research and application studies covering knowledge acquisition, knowledge base construction, and KG-based question answering.
There is no installable software in the primary use case. You clone the repository, navigate to the subdirectory relevant to your research area, and read the documents. The repository stores PDF and PPT files alongside Markdown notes and links to external analyses hosted on Baidu Mind Maps and other platforms.
The README opens by asking contributors who want to improve the project to get in touch. This signals a one-maintainer research notes project rather than a community software library, though it does list external resource links like NLP-Progress and Papers With Code for tracking the broader field.
Directory Structure: Topic-Organized Subdirectories
The top-level repository has around 15 subdirectories, each covering a specific research domain. The primary areas are:
Knowledge graph fundamentals (`知识图谱基础/`) holds foundational materials including presentations from the CCKS (China Knowledge Graph and Semantic Computing) conference series from 2013 through 2018. These are actual PDF and PPT files stored in the repository, not just links.
Knowledge base construction (`知识库构建/`) and knowledge storage (`知识存储/`) cover the technical process of building and querying KG backends, including graph database tools.
NLP research (`自然语言处理/`) and semantic computation (`语义计算/`) cover language model research, text preprocessing, and word representation methods.
The dialogue system area (`基于知识图谱的对话系统/`) is specifically about QA systems built on knowledge graphs, and it appears to include a project called 基于知识图谱的对话系统 as a standalone study.
Event graph (`事理图谱/`) covers event-centric knowledge graphs, a specialized research area that models causal and temporal relationships between events rather than entity relationships.
Chinese financial document processing (`中文金融文档智能处理/`) is a domain-specific section covering the chFinAnn dataset and financial event extraction research.
The MkDocs configuration at the root (`mkdocs.yml`) suggests the repository content is organized to build as a documentation website, though the README does not document how to build it.
Paper Analyses and Annotated Reading Lists
The most distinctive aspect of the repository is the annotated paper analyses section. Rather than linking to papers without context, the README provides mind-map analyses hosted on Baidu Mind Maps for several foundational papers and topics.
The KG and QA theory section includes annotated summaries for topics including KG survey papers, the challenges of knowledge graphs, the intersection of deep learning and knowledge graphs, CN-DBpedia (a Chinese knowledge extraction system), and the KBQA (knowledge base question answering) category.
The NLP paper analysis section covers the Transformer architecture (linking to Jay Alammar's illustrated Transformer), attention survey papers, BERT, ERNIE (two versions: Baidu's knowledge-enhanced ERNIE and Tsinghua's entity-enhanced version), and Google's T5.
These analysis links are a practical differentiator from a raw awesome-list. A researcher new to a topic gets a mental map of the paper before reading, which reduces the time to extract the key contribution.
The commercial applications section covers NLP use cases from Chinese industry players including Xiaomi, Xunfei (iFlytek), and a knowledge graph methodology presentation from Wenyin Technology. These are presented explicitly as educational references, not endorsements.
Tool and Dataset Reference Tables
Several sections in the README function as reference tables for practitioners who need to pick tools or datasets.
The mainstream QA and dialogue systems list includes three open-source systems: a Java-based QA system (QuestionAnsweringSystem), a medical knowledge graph QA system (QASystemOnMedicalKG in Python), and DeepPavlov, the open-source deep learning dialog library from MIPT.
The semantic platform list covers Chinese NLP platforms: Tencent WenzHi, iFlytek Open Semantic Platform, Boson NLP, and HIT LTP (Language Technology Platform from Harbin Institute of Technology).
The graph storage and query tools section and the visualization tools section are referenced in the README's table of contents but their content is not fully visible in the available portion.
The Chinese and English knowledge graph dataset list is a distinct section covering reference datasets for KG construction research. The conference calendar lists A-class and B-class academic venues: ACL, CVPR, ICML, IJCAI, EMNLP, CIKM, AAAI, SIGKDD, TKDE, and SIGIR, with their classification levels and domains.
What This Repository Does Not Cover
The repository's framing is knowledge graph-centric NLP research. It does not provide tutorials for applying LLMs to production tasks, benchmark results for modern models, or hands-on code implementations beyond the papers it summarizes.
The code-focused content mentioned in the directory structure comment, including NLP and KG algorithm implementations, text similarity experiments, and the SmartInteraction dialogue system project, is partially commented out in the README. The actual files exist in the subdirectories, but the README does not guide the reader through using them.
English-language content is sparse. The external links include a few English resources (NLP-Progress, Papers With Code, Jay Alammar's blog), but most of the original analysis, the CCKS conference presentations, and the commercial application references are in Chinese. A researcher working only in English will find limited value in the primary content.
The repository was created in 2019 and the README's analytical content reflects the research landscape of that period. LLM architectures beyond the transformer (GPT-style, instruction-tuned models, multimodal models) are referenced at a high level in the README's opening notes but are not developed into the same depth as the KG and early BERT-era content.
Comparison with English NLP Awesome-Lists
The most widely known English alternative is the keon/awesome-nlp repository, which maintains a categorized list of NLP resources, papers, and tools, updated by community contributions. Another common reference is the flairNLP/flair repository, which is a software library but includes extensive documentation of NLP tasks and benchmarks.
The difference with lihanghang/NLP-Knowledge-Graph is scope and language. Awesome-NLP lists tend to be link collections maintained by many contributors, with minimal original analysis. This repository adds mind-map paper summaries and conference presentations in Chinese, making it more useful for Chinese-speaking researchers who want curated analysis rather than a raw link index.
The repository is also more narrowly focused on knowledge graphs and QA systems than a general NLP awesome-list. A researcher specifically interested in KG construction, entity and relation extraction, or KG-based QA will find more directly relevant material here than in a broad NLP repository.
Maintenance and Licensing
The repository is not archived. The last push was on 2026-09-03, indicating ongoing maintenance. The repository was created in August 2019, so it has over seven years of accumulated content. There are no GitHub releases and no version numbering; content is updated by direct commits to the master branch.
The README includes a star history section, which notes the repository's growth as a community signal, but the README's instructions prohibit citing it as evidence of quality.
The license is MIT. All content in the repository, including the Markdown notes, PDF presentations, and external link collections, can be freely used, shared, and modified. The MIT license on a research notes repository is permissive and unusual; it means anyone can fork the content and republish it, with or without attribution, as long as the license text is preserved.
Contribution is invited: the README's opening note asks interested contributors to contact the author. There is no GitHub issue template or contributing guide in the repository.
Editorial conclusion
NLP-Knowledge-Graph is useful for Chinese-speaking researchers and engineers who want a single starting point for knowledge graph and NLP literature, curated with summaries and mind maps rather than raw paper lists. It is not a software package, and it does not cover recent LLM benchmarks or transformer architectures released after its last active update cycle. For English-speaking practitioners, the repository's value is limited because most of its linked content and analysis is in Chinese. The last push was on 2026-09-03, indicating the repository is still receiving updates, though the core content reflects a research period starting in 2019.
Frequently asked questions
Is NLP-Knowledge-Graph a software library you can install?
No. lihanghang/NLP-Knowledge-Graph is a research notes and reference collection, not an installable package. You clone the repository and navigate to the relevant subdirectory to read paper analyses, conference presentations, and tool comparison tables. The README does not document any install or run commands.
What topics does NLP-Knowledge-Graph cover?
The repository covers knowledge graph construction and querying, NLP paper analyses including BERT, ERNIE, and T5, dialogue and QA systems built on knowledge graphs, event graphs, Chinese financial document processing, and lists of academic conferences, semantic platforms, text preprocessing tools, and KG datasets.
Is NLP-Knowledge-Graph still being maintained?
The repository is not archived and the last push was on 2026-09-03. It has been receiving updates since its creation in August 2019. There are no formal releases; updates go directly to the master branch.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lihanghang-nlp-knowledge-graph)