rum&gin索引对比

RUM

RUM访问方法扩展了GIN的基础概念,使我们能够更快地执行全文搜索。 是GIN的下一代全新功能

GIN存在的限制

RUM让我们超越了GIN的哪些限制?

首先,<<tsvector>>数据类型不仅包含lexemes,而且还包含它们在文档中的位置信息。GIN索引并不存储这些信息。因此,GIN索引对搜索出现在短语搜索(全文搜索查询可以包含考虑词素之间距离的特殊运算符)操作的支持效率很低,并且必须访问原始数据进行重新检查。

其次,搜索系统通常根据相关性(不管那意味着什么)返回结果。 我们可以使用排序(ranking)函数<<ts_rank>>和<<ts_rank_cd>>来达到这个目的,但是它们必须对结果的每一行进行计算,这当然是很慢的。

近似地说,可以将RUM访问方法看作GIN,它额外存储位置信息,并可以按需要的顺序返回结果。

在RUM索引中,每个词素不只是引用表中的行:每个TID都提供了该词素在文档中出现的位置列表。

我们通过一组测试数据来对比以下gin和rum的差异

首先创建测试数据

kingbase=# drop table t1;

DROP TABLE

kingbase=# create table t1 as select name,short_desc from sys_settings;

SELECT 454

kingbase=# alter table t1 add column tsv tsvector;

ALTER TABLE

kingbase=# update t1 set tsv=to_tsvector(short_desc);

UPDATE 454

kingbase=# select * from t1 limit 1;

name | short_desc | tsv

--------------±---------------------±--------------------------------

row_security | Enable row security. | 'enable':1 'row':2 'security':3

(1 行记录)

然后我们查询 number和of 关键词紧挨着的查询结果

skingbase=# select * from t1 where to_tsvector(short_desc) @@ to_tsquery('number <-> of') limit 1;

name | short_desc | tsv

autovacuum_analyze_scale_factor | Number of tuple inserts, updates, or deletes prior to analyze as a fraction of reltuples. | 'a':12 'analyze':10 'as':11 'deletes':7 'fr action':13 'inserts':4 'number':1 'of':2,14 'or':6 'prior':8 'reltuples':15 'to':9 'tuple':3 'updates':5 (1 行记录)

kingbase=#

创建 gin 索引并查看执行计划

create index ind_t1_rum on t1 using gin(tsv);

kingbase=# explain analyze select short_desc from t1 where tsv @@ to_tsquery('number <-> of') order by tsv <=> to_tsquery('number <-> of') limit 1;

QUERY PLAN

Limit (cost=40.61...40.62 rows=1 width=59) (actual time=0.269...0.269 rows=1 loops=1)

-> Sort (cost=40.61...40.66 rows=17 width=59) (actual time=0.268...0.268 rows=1 loops=1)

Sort Key: ((tsv <=> to_tsquery('number <-> of'::text)))

Sort Method: top-N heapsort Memory: 25kB

-> Bitmap Heap Scan on t1 (cost=12.38...40.53 rows=17 width=59) (actual time=0.050...0.254 rows=50 loops=1)

Recheck Cond: (tsv @@ to_tsquery('number <-> of'::text))

Heap Blocks: exact=13

-> Bitmap Index Scan on ind_t1_rum (cost=0.00...12.38 rows=17 width=0) (actual time=0.032...0.033 rows=50 loops=1)

Index Cond: (tsv @@ to_tsquery('number <-> of'::text))

Planning Time: 0.291 ms

Execution Time: 0.289 ms

(11 行记录)

kingbase=# explain analyze select short_desc from t1 where tsv @@ to_tsquery('number <-> of') ;

QUERY PLAN

Bitmap Heap Scan on t1 (cost=12.38...36.24 rows=17 width=55) (actual time=0.037...0.142 rows=50 loops=1)

Recheck Cond: (tsv @@ to_tsquery('number <-> of'::text))

Heap Blocks: exact=13

-> Bitmap Index Scan on ind_t1_rum (cost=0.00...12.38 rows=17 width=0) (actual time=0.025...0.025 rows=50 loops=1)

Index Cond: (tsv @@ to_tsquery('number <-> of'::text))

Planning Time: 0.144 ms

Execution Time: 0.161 ms

创建rum 索引并查看执行计划

kingbase=# create index ind_t1_rum on t1 using rum(tsv);

CREATE INDEX

kingbase=# explain analyze select short_desc from t1 where tsv @@ to_tsquery('number <-> of') order by tsv <=> to_tsquery('number <-> of') limit 1;

QUERY PLAN

Limit (cost=8.25...11.58 rows=1 width=59) (actual time=0.081...0.082 rows=1 loops=1)

-> Index Scan using ind_t1_rum on t1 (cost=8.25...64.84 rows=17 width=59) (actual time=0.080...0.081 rows=1 loops=1)

Index Cond: (tsv @@ to_tsquery('number <-> of'::text))

Order By: (tsv <=> to_tsquery('number <-> of'::text))

Planning Time: 0.395 ms

Execution Time: 0.099 ms

(6 行记录)

kingbase=# explain analyze select short_desc from t1 where tsv @@ to_tsquery('number <-> of') ;

QUERY PLAN

Bitmap Heap Scan on t1 (cost=12.38...36.24 rows=17 width=55) (actual time=0.063...0.086 rows=50 loops=1)

Recheck Cond: (tsv @@ to_tsquery('number <-> of'::text))

Heap Blocks: exact=13

-> Bitmap Index Scan on ind_t1_rum (cost=0.00...12.38 rows=17 width=0) (actual time=0.037...0.037 rows=50 loops=1)

Index Cond: (tsv @@ to_tsquery('number <-> of'::text))

Planning Time: 0.180 ms

Execution Time: 0.110 ms

(7 行记录)

kingbase=#

通过对比执行计划可以发现,RUM索引创建之后对数据进行了排序,所以在通过limit进行查询时,不需要进行额外的sort操作,结果返回速度更快。

原因是使用RUM index,使用简单的索引扫描执行查询:不需要查看额外的文档,也不需要单独排序:(因为是精确查询,所以不需要重新检查;因为rum索引会根据"附加信息"------距离来组织索引项,所以不需要排序)

但是因为RUM比GIN储存更多的信息,所以它的尺寸必须更大。我们比较不同索引的大小;让我们把RUM加到这张表上:

RUM | GIN | GiST | btree

457MB | 179MB | 125MB | 546MB

正如我们所看到的,规模显著增长,这就是快速搜索的成本。

相关推荐
yychen_java1 小时前
第六篇:Spring AI 实战:将 Java 业务接口封装成企业级 MCP Server
java·人工智能·spring
wefg12 小时前
【Redis】初识 Redis
数据库·redis·缓存
_upupup2 小时前
异常(C++)
java·开发语言·jvm
g10565591392 小时前
公有云_云运维服务
java·运维·服务器
霸道流氓气质2 小时前
ComfyUI 图像生成工作流完全指南:从节点编排到Java生产级图像生产实战
java·开发语言·人工智能
再写一行代码就下班2 小时前
linux sh脚本在windows修改导致无法使用解决方式
java·linux·centos
m0_587383003 小时前
24小时自助健身房软硬件解决方案实战指南:从架构设计到部署实施
java·spring·小程序·架构·需求分析
陆卿之3 小时前
Java对接DeepSeek
java·开发语言·人工智能
庄园特聘拆椅狂魔3 小时前
Java 后端转全栈的第一课:从前端项目搭建到技术选
java·开发语言·前端