RUM
RUM访问方法扩展了GIN的基础概念,使我们能够更快地执行全文搜索。 是GIN的下一代全新功能
GIN存在的限制
RUM让我们超越了GIN的哪些限制?
首先,<<tsvector>>数据类型不仅包含lexemes,而且还包含它们在文档中的位置信息。GIN索引并不存储这些信息。因此,GIN索引对搜索出现在短语搜索(全文搜索查询可以包含考虑词素之间距离的特殊运算符)操作的支持效率很低,并且必须访问原始数据进行重新检查。
其次,搜索系统通常根据相关性(不管那意味着什么)返回结果。 我们可以使用排序(ranking)函数<<ts_rank>>和<<ts_rank_cd>>来达到这个目的,但是它们必须对结果的每一行进行计算,这当然是很慢的。
近似地说,可以将RUM访问方法看作GIN,它额外存储位置信息,并可以按需要的顺序返回结果。
在RUM索引中,每个词素不只是引用表中的行:每个TID都提供了该词素在文档中出现的位置列表。
我们通过一组测试数据来对比以下gin和rum的差异
首先创建测试数据
kingbase=# drop table t1;
DROP TABLE
kingbase=# create table t1 as select name,short_desc from sys_settings;
SELECT 454
kingbase=# alter table t1 add column tsv tsvector;
ALTER TABLE
kingbase=# update t1 set tsv=to_tsvector(short_desc);
UPDATE 454
kingbase=# select * from t1 limit 1;
name | short_desc | tsv
--------------±---------------------±--------------------------------
row_security | Enable row security. | 'enable':1 'row':2 'security':3
(1 行记录)
然后我们查询 number和of 关键词紧挨着的查询结果
skingbase=# select * from t1 where to_tsvector(short_desc) @@ to_tsquery('number <-> of') limit 1;
name | short_desc | tsv
autovacuum_analyze_scale_factor | Number of tuple inserts, updates, or deletes prior to analyze as a fraction of reltuples. | 'a':12 'analyze':10 'as':11 'deletes':7 'fr action':13 'inserts':4 'number':1 'of':2,14 'or':6 'prior':8 'reltuples':15 'to':9 'tuple':3 'updates':5 (1 行记录)
kingbase=#
创建 gin 索引并查看执行计划
create index ind_t1_rum on t1 using gin(tsv);
kingbase=# explain analyze select short_desc from t1 where tsv @@ to_tsquery('number <-> of') order by tsv <=> to_tsquery('number <-> of') limit 1;
QUERY PLAN
Limit (cost=40.61...40.62 rows=1 width=59) (actual time=0.269...0.269 rows=1 loops=1)
-> Sort (cost=40.61...40.66 rows=17 width=59) (actual time=0.268...0.268 rows=1 loops=1)
Sort Key: ((tsv <=> to_tsquery('number <-> of'::text)))
Sort Method: top-N heapsort Memory: 25kB
-> Bitmap Heap Scan on t1 (cost=12.38...40.53 rows=17 width=59) (actual time=0.050...0.254 rows=50 loops=1)
Recheck Cond: (tsv @@ to_tsquery('number <-> of'::text))
Heap Blocks: exact=13
-> Bitmap Index Scan on ind_t1_rum (cost=0.00...12.38 rows=17 width=0) (actual time=0.032...0.033 rows=50 loops=1)
Index Cond: (tsv @@ to_tsquery('number <-> of'::text))
Planning Time: 0.291 ms
Execution Time: 0.289 ms
(11 行记录)
kingbase=# explain analyze select short_desc from t1 where tsv @@ to_tsquery('number <-> of') ;
QUERY PLAN
Bitmap Heap Scan on t1 (cost=12.38...36.24 rows=17 width=55) (actual time=0.037...0.142 rows=50 loops=1)
Recheck Cond: (tsv @@ to_tsquery('number <-> of'::text))
Heap Blocks: exact=13
-> Bitmap Index Scan on ind_t1_rum (cost=0.00...12.38 rows=17 width=0) (actual time=0.025...0.025 rows=50 loops=1)
Index Cond: (tsv @@ to_tsquery('number <-> of'::text))
Planning Time: 0.144 ms
Execution Time: 0.161 ms
创建rum 索引并查看执行计划
kingbase=# create index ind_t1_rum on t1 using rum(tsv);
CREATE INDEX
kingbase=# explain analyze select short_desc from t1 where tsv @@ to_tsquery('number <-> of') order by tsv <=> to_tsquery('number <-> of') limit 1;
QUERY PLAN
Limit (cost=8.25...11.58 rows=1 width=59) (actual time=0.081...0.082 rows=1 loops=1)
-> Index Scan using ind_t1_rum on t1 (cost=8.25...64.84 rows=17 width=59) (actual time=0.080...0.081 rows=1 loops=1)
Index Cond: (tsv @@ to_tsquery('number <-> of'::text))
Order By: (tsv <=> to_tsquery('number <-> of'::text))
Planning Time: 0.395 ms
Execution Time: 0.099 ms
(6 行记录)
kingbase=# explain analyze select short_desc from t1 where tsv @@ to_tsquery('number <-> of') ;
QUERY PLAN
Bitmap Heap Scan on t1 (cost=12.38...36.24 rows=17 width=55) (actual time=0.063...0.086 rows=50 loops=1)
Recheck Cond: (tsv @@ to_tsquery('number <-> of'::text))
Heap Blocks: exact=13
-> Bitmap Index Scan on ind_t1_rum (cost=0.00...12.38 rows=17 width=0) (actual time=0.037...0.037 rows=50 loops=1)
Index Cond: (tsv @@ to_tsquery('number <-> of'::text))
Planning Time: 0.180 ms
Execution Time: 0.110 ms
(7 行记录)
kingbase=#
通过对比执行计划可以发现,RUM索引创建之后对数据进行了排序,所以在通过limit进行查询时,不需要进行额外的sort操作,结果返回速度更快。
原因是使用RUM index,使用简单的索引扫描执行查询:不需要查看额外的文档,也不需要单独排序:(因为是精确查询,所以不需要重新检查;因为rum索引会根据"附加信息"------距离来组织索引项,所以不需要排序)
但是因为RUM比GIN储存更多的信息,所以它的尺寸必须更大。我们比较不同索引的大小;让我们把RUM加到这张表上:
RUM | GIN | GiST | btree
457MB | 179MB | 125MB | 546MB
正如我们所看到的,规模显著增长,这就是快速搜索的成本。