声明:从书中蒸馏知识"目前不是一个单一、统一命名的研究领域,而是分散在 Textbook Knowledge Extraction、Textbook Knowledge Graph、LLM-based Knowledge Extraction、Knowledge Distillation、Educational LLM 等几个方向。
《Knowledge Graphs for Textbooks: Extraction and Completion Techniques》《教材知识图谱:抽取与补全技术》
原文摘要:
This paper aims to apply knowledge graph construction techniques to textbooks, explicitly focusing on the challenge of the absence of domain-specific schema for each textbook. Various entity and relation extraction models are utilized to capture logical and semantic information related to the textbook's topic. These models include a Text-Encoding-Initiative (TEI) model to extract hierarchical concepts, spaCy Natural Language Processing (NLP), and Google Cloud Natural Language to extract semantic information from the main textual content. The study includes a case study on a cloud computing textbook, where each approach is evaluated and analyzed. Ultimately, the goal is to create knowledge graphs of textbooks, enabling the completion task of predicting missing entities or relations in a low-dimensional space.
重点解决:教材缺少领域专属本体 Schema 这一难题。(也就是说:没有一套预先定义好、专门适配这本教材学科的「实体类型、关系类型、约束规则」。 Schema = 知识图谱的数据库表结构 + 概念词典;本体 Ontology = 带有语义约束的 Schema。)
Ontology和Schema有什么关系: 放在这篇教材 KG 论文语境下:Schema 是结构骨架;Ontology 是带语义约束、概念定义的完整知识体系。Ontology 包含 Schema,Schema 是 Ontology 里的 "数据结构部分"。

**核心模块:**知识抽取(从教材文本 / 图表提取实体、关系、三元组)+知识补全(KG Completion,补隐式关系、缺失实体,修正逻辑冲突)
关键点:这篇论文没有使用 LLM(2023 AIKE,早于大模型热潮),用传统 NLP+KGE 嵌入做教材 KG;和现在 Textbook2KG、K12-KGraph 路线不一样。
什么是显式三元组,什么是隐式三元组?
回答:
显式三元组:文本明确陈述,句子里直接包含这个事实,不需要脑补、推导、猜测。
获取方式:知识抽取
例如:原文句子:加速度是速度的变化量与发生这一变化所用时间的比值。 可以直接抽出显式三元组: (加速度,定义为,速度的变化量与时间的比值)
特点:有原文句子作为直接证据 ;不需要逻辑推理;抽取模型 / LLM 只需要读懂句子,就能提取;可信度高,不属于知识补全的结果。
隐式三元组:原文没有直接写,需要推理才能得到的关系,这就是知识图谱补全(KG Completion)要干的。
获取方式:知识补全(需要逻辑,容易产生幻觉错误)
还是上面那句加速度定义:
原文只定义加速度,没有写:学习加速度需要先掌握速度。
但是从教学逻辑我们可以推断: (加速度,前置知识,速度)