Model or dataset
zjunlp/DeepKE avatar
zjunlp/DeepKE

DeepKE: the knowledge extraction toolkit behind a decade of ZJUNLP models

[EMNLP 2022] An Open Toolkit for Knowledge Graph Extraction and Construction

4,486 stars750 forksPythonMIT

At a glance

What is it?
DeepKE is ZJUNLP's MIT-licensed Python toolkit for knowledge graph construction, published at EMNLP 2022, supporting cnSchema, low-resource, document-level and multimodal scenarios for entity, relation, attribute and event extraction. It bundles published models from LightNER and W2NER to KnowPrompt, ASP, PRGC and PURE, extends to LLM-based extraction through DeepKE-LLM and the OneKE model, and ships MCP tools so language models can call it.
Who is it for?
Use DeepKE when building knowledge graphs from text with supervised models, whether named entity recognition, relation extraction, relational triples or events, and when published research implementations with consistent configuration matter, since its example tree is organized by task with Hydra configuration. For LLM-first extraction, follow the DeepKE-LLM and OneKE paths instead of the classic training pipelines.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 79 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Four scenarios, three or four tasks

DeepKE is a knowledge extraction toolkit for knowledge graph construction supporting cnSchema, low-resource, document-level and multimodal scenarios for entity, relation and attribute extraction, and the table of contents adds event extraction as the fourth function. Each scenario name maps to a research problem the toolkit implements, cnSchema to the ready-to-use Chinese schema models, low-resource to few-shot methods, document-level to extraction across sentences rather than within them, and multimodal to extraction combining text with images. The toolkit's own framing is a deep learning based knowledge extraction toolkit for knowledge graph construction, and the introduction pairs it with documentation, an online demo, the paper, slides and a poster, the full research-artifact set a EMNLP publication carries.

The model roster, by paper

The quick start names the supervised models with their venues. Named entity recognition offers LightNER from COLING 2022 in the few-shot directory and W2NER from AAAI 2022 under standard. Relation extraction includes KnowPrompt from WWW 2022 in the few-shot path. Relational triple extraction carries three generations, ASP from EMNLP 2022, PRGC from ACL 2021 and PURE from NAACL 2021. April 2023 added CP-NER from IJCAI 2023 for cross-type NER. The example tree mirrors this exactly, directories for ner, re, triple, ae, ee and llm, so a researcher reproducing a paper finds its implementation by venue year, the organization decision that makes the toolkit usable as a research baseline collection rather than only a product. The few-shot and standard split inside each task directory carries its own meaning, few-shot methods target the low-resource scenario with prompt-based or lightweight adaptation, while the standard directories hold the conventional supervised training pipelines, so a user with abundant labels and a user with fifty examples take different doors into the same task.

DeepKE-LLM and the OneKE thread

The LLM path has its own lineage in the news section. February 2023 added GPT-3 with in-context learning based on EasyInstruct plus data generation. June 2023 extended DeepKE-LLM to knowledge extraction with KnowLM, ChatGLM, the LLaMA series and the GPT series. September 2023 released the InstructIE bilingual instruction dataset for instruction-based knowledge graph construction. February 2024 shipped IEPile, a 0.32 billion token bilingual IE instruction dataset, with baichuan2 and llama2 13B LoRA models trained on it. April 2024 released OneKE, a schema-based bilingual extraction model built on Chinese-Alpaca-2-13B, and December 2024 open sourced the OneKE framework for multi-agent knowledge extraction. The README routes LLM-curious users to DeepKE-LLM and OneKE before the classic paths, reading the field's shift directly into the document's order.

MCP tools: the toolkit as an LLM's function

The June 2025 news added MCP service tools to DeepKE, enabling knowledge extraction through large language models as tool callers for lightweight models, hosted on ModelScope's MCP server registry, and the repository carries an mcp-tools directory beside its top-level entries. The direction reverses the DeepKE-LLM relationship, there a language model performs extraction guided by DeepKE's data, here the classical models become callable tools that an agent invokes, the architecture where a heavyweight reasoning model delegates named entity recognition to a small trained specialist. The addition also aligns with the CCF ODTC open source incentive program selection announced in July 2026, the most recent news entry, recognizing the toolkit's continued service as infrastructure.

cnSchema models, ready to run

Beyond training your own, the toolkit releases off-the-shelf models at DeepKE-cnSchema, and the repository maintains separate readme files for cnSchema in English and Chinese, giving that path its own documentation weight. cnSchema is the Chinese knowledge schema the models target, so the released models extract entities and relations conforming to a shared vocabulary without the user defining one, the lowest-friction entry into Chinese knowledge graph construction the toolkit offers. The prediction demo and the online demo at deepke.zjukg.cn let a user try extraction before any installation, and the Colab notebook linked at the top provides a hosted environment for the first run, the three zero-setup paths beside the local install.

Requirements, pinned for the classic stack

The requirements file pins the classic training stack precisely, torch between 1.5 and 1.11, hydra-core at 1.0.6, transformers at 4.26.0, alongside jieba for Chinese tokenization, seqeval for NER metrics, pytorch-crf for sequence labeling, scikit-learn, wandb and tensorboard for tracking, and notably openai at 0.28, the older completions-era API surface the LLM examples grew up with. The Hydra dependency is the configuration backbone, each example's config files drive the experiments, and the platform guidance in the introduction is practical, Linux is recommended, and on Windows file paths need double backslashes. The introduction also anticipates the most common installation failure, if HuggingFace is inaccessible, consider wisemodel or modescape as model mirrors, advice written for users behind network restrictions. The annotation instructions and weak-supervision automatic labelling added in November 2022 cover the step before training, generating labeled data for entity and relation extraction, which in practice decides whether a knowledge graph project ships or stalls, since models are rarely the bottleneck that data is.

A toolkit with a family tree

The readme's closing sections trace the ecosystem, reading materials, related toolkit, citation, contributors and a section on other knowledge extraction open source projects from the same lab. The citation carries a CITATION.cff file at the repository root for tooling, and the related projects section connects DeepKE to ZJUNLP's broader output. The package itself is deepke on PyPI at version 2.2.7, with releases tagged through 2023 while the main branch continued, last pushed 2026-07-13. The MIT license, bilingual documentation, and the demo-plus-paper-plus-slides-plus-poster set complete the picture of a university toolkit maintained as a public research instrument, where the measure of success is downstream papers and built graphs rather than deployment counts.

Editorial conclusion

Use DeepKE when building knowledge graphs from text with supervised models, whether named entity recognition, relation extraction, relational triples or events, and when published research implementations with consistent configuration matter, since its example tree is organized by task with Hydra configuration. For LLM-first extraction, follow the DeepKE-LLM and OneKE paths instead of the classic training pipelines. Before adopting, plan for Linux, use double backslashes in paths on Windows, note the requirements pin older torch and transformers versions for the classic stack, check the tips section when installation fails, and consider the off-the-shelf cnSchema models before training from scratch.

Frequently asked questions

What is knowledge extraction?

Knowledge extraction is deriving structured information from text for knowledge graph construction, and DeepKE implements it across four functions, named entity recognition, relation extraction, attribute extraction and event extraction, supporting cnSchema, low-resource, document-level and multimodal scenarios.

What is the difference between DeepKE and OneKE?

DeepKE is the classic supervised toolkit with per-task models like LightNER, W2NER and PRGC trained through its Hydra-configured pipelines, while OneKE is the LLM-based line, a schema-based bilingual extraction model on Chinese-Alpaca-2-13B grown into an open-source multi-agent knowledge extraction framework, with DeepKE-LLM as the bridge supporting GPT, ChatGLM, LLaMA and KnowLM models.

How do you start with DeepKE without training models?

Use the released off-the-shelf DeepKE-cnSchema models for extraction conforming to the cnSchema vocabulary, try the online demo at deepke.zjukg.cn or the Colab notebook before installing anything, and for LLM-based extraction follow the DeepKE-LLM examples or the OneKE model rather than the classic training pipelines.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. zjunlp/DeepKE on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/zjunlp-deepke.svg)](https://hysenlabs.com/projects/zjunlp-deepke)