【从零开始学习 RAG 】01:LLamaIndex 基本概念

【从零开始学习 RAG 】01:LLamaIndex 基本概念

从0开始的llamaindex学习,不能只学langchain一个框架:

今天从 llamaindex 表达数据的基本类型 Documents 和 Nodes 开始

0. 环境配置

我主要使用 LM Studio 本地运行LLM。LM Studio 最大的优势在于可以适配 OpenAI API 接口,这样就可以很方便在本地调试成功后,后续只需要轻微的修改就可以切换到 OpenAI API 了。

LM Studio 在启动 Server 的时候,很贴心地提供了调用 API 的代码,以下实践都是基于此进行修改的。

本次测试使用到的模型是 mistralai/Mixtral-8x7B-Instruct-v0.1 · Hugging Face。在实际使用中接近 GPT3.5。作为本地测试模型性能完全足够了。

1. Documents 和 Nodes

Documents

Documents 是 llamaindex 中描述文件的基本类型,它是"任意类型文件"的存储容器。这些文件可以是"pdf文本"、"图片",甚至是"向量数据"。

Documents 存储着 2 个核心"元数据":

  • metadata​ - a dictionary of annotations that can be appended to the text.

    • 描述 Document 的基本信息
  • relationships​ - a dictionary containing relationships to other Documents/Nodes.

    • 文件的"关系",一般用于表示多个 Document 之间的关系上
    • 多个 Document 节点可以根据 Relationship 组合成"网状结构"

测试 Documents 的代码:

py 复制代码
# https://docs.llamaindex.ai/en/stable/module_guides/loading/documents_and_nodes/root.html
# Document is the important type in the LLaMaIndex

from llama_index import SimpleDirectoryReader

# load data from file
document = SimpleDirectoryReader(
    input_files=["./data/king.dreamspeech.excerpts.pdf"]
).load_data()

# Look into the type of "Document"
print(type(document), '\n')
print(len(document), '\n')
print(type(document[0]), '\n')
print(document[0])
  • 通过 SimpleDirectoryReader 读取文本文件,作为 Document 的赋值

Nodes

以上仅仅是"加载好数据"而已,但是对于 RAG 系统存储数据的形式来说,Document 并不是最终的结果。在向量数据库中,数据都以"分块"的形式进行存储,英文翻译为"chunk"。

不过在 llamaindex 中,这一概念解释为"Node"。

"Node"在 llamaindex 中是 Document 的一个"组成部分"。单一的 Document 会被算法切分为多个 Node 进行存储。

同样,对于 Node 来说,它也拥有 metadata 和 relationship 等元数据。也意味着多个 node 可以组合成"网状信息结构"。

测试代码:

py 复制代码
# https://docs.llamaindex.ai/en/stable/module_guides/loading/documents_and_nodes/root.html
from llama_index import SimpleDirectoryReader
from llama_index.node_parser import SentenceSplitter

# load data from file
document = SimpleDirectoryReader(
    input_files=["./data/king.dreamspeech.excerpts.pdf"]
).load_data()

# get nodes from document
parser = SentenceSplitter(
    chunk_size=100,
    chunk_overlap=10,
)
nodes = parser.get_nodes_from_documents(document)

# Look into the type of "Node"
print(type(nodes), '\n')
print(len(nodes), '\n')
print(type(nodes[0]), '\n')
print(nodes[0])

for i in range(len(nodes)):
    print(nodes[i], '\n')

一部分的数据结果:

py 复制代码
Node ID: 55034aee-9a47-4458-b2a8-2ede141552fd
Text: No, no, we are not satisfied, and we will not be satisfied until
justice rolls down like wat ers and  righteousness like a mighty
stream.  . . .  I say to you today, my friends, though, even though we
face the difficulties of today and tomorrow, I still

Node ID: 333dbfac-dfae-442b-99bf-d2b0d71cd59d
Text: ©2014 The Gilder Lehrman Institute of American History
www.gilderlehrman.org  have a dream. It is a dream deeply rooted in
the American dream. I have a dream that one day this  nation will rise
up, live out the true meaning of its creed: "We hold these truths to
be  self- evident,

项目源代码

相关推荐
TheBestRucy20 分钟前
RAG知识库问答系统落地:从向量检索到上下文增强的全链路实践
人工智能·python·langchain·aigc·交互
智行合一科技44 分钟前
WAIC释放关键信号:合规能力正成为AIGC营销的核心竞争壁垒
aigc
爱听歌的周童鞋3 小时前
霹雳吧啦Wz | AIGC | 图像生成篇 | DDPM介绍与公式推导 | 笔记 | (一) 前向加噪 & 优化目标引入
aigc·diffusion model·ddpm·reverse process·image generate·forward process
longxibo3 小时前
第 15 章 政务/制造业落地案例
人工智能·深度学习·aigc·政务
卡卡罗特AI4 小时前
AI编程入门教程02-LLM发展历程,AI 御四家十年风云:OpenAI 分裂、Anthropic 出走、谷歌掉队、马斯克上桌
aigc·openai·ai编程
Am-Chestnuts4 小时前
AI长回答批量导出PDF与长图:多轮内容整理和免费Markdown备份
人工智能·pdf·aigc
leeyi4 小时前
RAG 流水线设计:Eino 的 Loader → Transformer → Indexer → Retriever(第60篇-E46)
aigc·agent·ai编程
Token炼金师4 小时前
框架的擂台:LangChain、LlamaIndex、Dify、AutoGen 与 LangGraph —— 应用框架选型五局
人工智能·深度学习·llm
wenb1n4 小时前
【 LLM】Agent Planning 完全指南:8 种纯 LLM 范式 + 8 种混合规划模式详解(一)
llm
bonechips4 小时前
LLM 的"严谨"与"胡说":temperature、Top K + LangChain 工作流
langchain·llm