使用 Elastic Stack 丰富邮政地址信息 —— 第二部分

作者:来自 Elastic

在上一篇文章中,我们介绍了如何将来自 BANO 项目的数据导入 Elasticsearch 并建立索引。现在,我们已经拥有包含法国所有邮政地址信息的索引。

接下来,让我们看看如何利用这个数据集实现更多功能。

搜索地址

很好。现在,我们能否使用搜索引擎来查询这些地址呢?

bash 复制代码
`

1.  GET .bano/_search?search_type=dfs_query_then_fetch
2.  {
3.    "size": 1,
4.    "query": {
5.      "bool": {
6.        "should": [
7.          {
8.            "match": {
9.              "address.number": "23"
10.            }
11.          },
12.          {
13.            "match": {
14.              "address.street_name": "r verdiere"
15.            }
16.          },
17.          {
18.            "match": {
19.              "address.city": "rochelle"
20.            }
21.          }

23.        ]
24.      }
25.    }
26.  }

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

查询返回的结果如下:

bash 复制代码
`

1.  {
2.    "took": 170,
3.    "timed_out": false,
4.    "_shards": {
5.      "total": 99,
6.      "successful": 99,
7.      "skipped": 0,
8.      "failed": 0
9.    },
10.    "hits": {
11.      "total": 10380977,
12.      "max_score": 23.681055,
13.      "hits": [
14.        {
15.          "_index": ".bano-17",
16.          "_type": "doc",
17.          "_id": "173008250H-23",
18.          "_score": 23.681055,
19.          "_source": {
20.            "address": {
21.              "zipcode": "17000",
22.              "number": "23",
23.              "city": "La Rochelle",
24.              "street_name": "Rue Verdière"
25.            },
26.            "location": {
27.              "lon": -1.155167,
28.              "lat": 46.157353
29.            },
30.            "id": "173008250H-23",
31.            "source": "C+O",
32.            "region": "17"
33.          }
34.        }
35.      ]
36.    }
37.  }

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

在我的机器上,这次查询耗时 170ms。不过,如果我们事先知道地址所属的省份编号,就可以进一步提高查询速度。

因此,我们可以通过减少查询需要访问的分片数量来优化搜索性能:

ini 复制代码
`

1.  GET .bano-17/_search?search_type=dfs_query_then_fetch
2.  {
3.    // Same query
4.  }

`AI写代码

现在,查询返回了相同的结果,但速度快了很多:

bash 复制代码
`

1.  {
2.    "took": 6,
3.    "timed_out": false,
4.    "_shards": {
5.      "total": 1,
6.      "successful": 1,
7.      "skipped": 0,
8.      "failed": 0
9.    },
10.    "hits": {
11.      "total": 212872,
12.      "max_score": 18.955963,
13.      "hits": [
14.        {
15.          "_index": ".bano-17",
16.          "_type": "doc",
17.          "_id": "173008250H-23",
18.          "_score": 18.955963,
19.          "_source": {
20.            "address": {
21.              "zipcode": "17000",
22.              "number": "23",
23.              "city": "La Rochelle",
24.              "street_name": "Rue Verdière"
25.            },
26.            "location": {
27.              "lon": -1.155167,
28.              "lat": 46.157353
29.            },
30.            "id": "173008250H-23",
31.            "source": "C+O",
32.            "region": "17"
33.          }
34.        }
35.      ]
36.    }
37.  }

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

请注意,这次查询找到了 212,872 条地址记录,但相关性最高的结果正是我们要查找的地址。

没错,相关性(Relevance)是搜索引擎最重要的核心能力之一。

根据地理坐标搜索

我们知道,还可以通过地理坐标( Geo Point)来搜索地址。

首先,让我们采用一种最简单、最直接的方法。

我们搜索数据集中的所有地理位置点,但要求根据它们与输入坐标之间的距离进行排序。

这意味着,返回的第一条结果就是距离输入坐标最近的地址:

bash 复制代码
`

1.  GET .bano/_search
2.  {
3.    "size": 1, 
4.    "sort": [
5.      {
6.        "_geo_distance": {
7.          "location": {
8.            "lat": 46.15735,
9.            "lon": -1.1551
10.          }
11.        }
12.      }
13.    ]
14.  }

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

查询结果如下:

json 复制代码
`

1.  {
2.    "took": 403,
3.    "timed_out": false,
4.    "_shards": {
5.      "total": 99,
6.      "successful": 99,
7.      "skipped": 0,
8.      "failed": 0
9.    },
10.    "hits": {
11.      "total": 16402853,
12.      "max_score": null,
13.      "hits": [
14.        {
15.          "_index": ".bano-17",
16.          "_type": "doc",
17.          "_id": "173008250H-23",
18.          "_score": null,
19.          "_source": {
20.            "address": {
21.              "zipcode": "17000",
22.              "number": "23",
23.              "city": "La Rochelle",
24.              "street_name": "Rue Verdière"
25.            },
26.            "location": {
27.              "lon": -1.155167,
28.              "lat": 46.157353
29.            },
30.            "id": "173008250H-23",
31.            "source": "C+O",
32.            "region": "17"
33.          },
34.          "sort": [
35.            5.176690615711886
36.          ]
37.        }
38.      ]
39.    }
40.  }

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

这里有几点值得注意。首先,我们在查询中指定了以下条件:

markdown 复制代码
`

1.  {
2.    "lat": 46.15735,
3.    "lon": -1.1551  
4.  }

`AI写代码

我们得到的是另一个地理坐标点。当然,BANO 并没有记录每一厘米的地理位置信息,而只是存储了各个地址对应的坐标:

markdown 复制代码
`

1.  {
2.    "lat": 46.157353,
3.    "lon": -1.155167
4.  }

`AI写代码

第二点需要注意的是,参与距离排序的命中文档总数为 16,402,853,也就是整个数据集中的所有地址记录。

这直接影响了查询的响应时间:403 毫秒(403ms)。

我们可以合理地假设,当我们根据地理坐标查找地址时,通常能够在该坐标周围大约 2 公里的范围内找到匹配的地址。

因此,我们不必对数据集中的所有地理坐标点进行排序,而是可以先根据 地理距离 对数据集进行过滤,再对符合条件的结果进行排序:

markdown 复制代码
`

1.  GET .bano/_search
2.  {
3.    "size": 1,
4.    "query": {
5.      "bool": {
6.        "filter": {
7.          "geo_distance": {
8.            "distance": "1km",
9.            "location": {
10.              "lat": 46.15735,
11.              "lon": -1.1551
12.            }
13.          }
14.        }
15.      }
16.    },
17.    "sort": [
18.      {
19.        "_geo_distance": {
20.          "location": {
21.            "lat": 46.15735,
22.            "lon": -1.1551
23.          }
24.        }
25.      }
26.    ]
27.  }

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

这次查询返回了相同的地理坐标点,但响应头中的信息有所不同:

json 复制代码
`

1.  {
2.    "took": 45,
3.    "timed_out": false,
4.    "_shards": {
5.      "total": 99,
6.      "successful": 99,
7.      "skipped": 0,
8.      "failed": 0
9.    },
10.    "hits": {
11.      "total": 4467,
12.      "max_score": null,
13.      "hits": [ /* ... */ ]
14.    }
15.  }

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

查询耗时仅为 45 毫秒(45ms),因为我们将需要参与排序的地理坐标点数量大幅减少到了 4,467 个。

如果我们事先知道地址所属的省份编号,还可以进一步优化查询性能:

arduino 复制代码
`

1.  GET .bano-17/_search
2.  {
3.    // Same query
4.  }

`AI写代码

现在,我们已经获得了相当理想的查询响应时间。

需要注意的是,文件系统缓存(Filesystem Cache) 在这里也发挥了非常重要的作用:

bash 复制代码
`

1.  {
2.    "took": 2,
3.    "timed_out": false,
4.    "_shards": {
5.      "total": 1,
6.      "successful": 1,
7.      "skipped": 0,
8.      "failed": 0
9.    }
10.  }

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

现在,我们已经知道如何查询这些数据了。接下来,让我们利用这些查询能力,通过 Logstash 读取现有数据集,并对其进行数据增强。

逐步构建 Logstash 数据增强管道

如果你一直在关注我的博客,就会知道,我通常会从一个基础的 Pipeline(数据处理管道)配置开始,如下所示(配置文件名为 bano-enrich.conf):

markdown 复制代码
`

1.  input { 
2.    stdin { } 
3.  }

5.  filter {
6.  }

8.  output {
9.    stdout { codec => rubydebug }
10.  }

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

使用起来非常简单,只需运行以下命令:

bash 复制代码
`head -1 mydata | bin/logstash -f bano-enrich.conf`AI写代码

到目前为止,一切都很顺利。但问题出在哪里呢?

每当你修改配置文件并希望测试修改后的效果时,都需要重新执行相同的命令行。

这本身并不是什么大问题,毕竟你的终端很可能保存了命令历史记录。😄

真正的问题在于 Logstash 的启动时间,包括 JVM 的启动以及 Logstash 自身的初始化过程。

在我的笔记本电脑上,整个启动过程大约需要 20 秒。

在我看来,这种开发体验并不算友好。

接下来,让我分享一个小技巧,帮助改善这个问题:改用 http-input-plugin(HTTP 输入 插件 ):

markdown 复制代码
`

1.  input {
2.    http { }
3.  }

`AI写代码

这会启动一个监听 8080 端口的 HTTP 服务器,你可以通过运行类似下面的命令来使用它:

vbnet 复制代码
`

1.  curl -XPOST "localhost:8080" -H "Content-Type: application/json" -d '{
2.    "test_case": "Address with text",
3.    "name": "Joe Smith",
4.    "address": {
5.      "number": "23",
6.      "street_name": "r verdiere",
7.      "city": "rochelle",
8.      "country": "France"
9.    }
10.  }'

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

这两种方式有什么区别呢?

由于我们不再使用标准输入(stdin),因此可以让 Logstash 在每次保存配置文件的新版本时,自动热重载(Hot Reload)数据处理管道,而无需重新启动 Logstash:

bash 复制代码
`bin/logstash -r -f bano-enrich.conf`AI写代码

之后,当你更新 bano-enrich.conf 配置文件时,就可以在 Logstash 日志中看到以下信息:

ini 复制代码
`

1.  [2018-03-24T11:06:17,680][INFO ][logstash.pipelineaction.reload] Reloading pipeline {"pipeline.id"=>:main}
2.  [2018-03-24T11:06:18,007][INFO ][logstash.pipeline        ] Pipeline has terminated {:pipeline_id=>"main", :thread=>"#<Thread:0x457a565f run>"}
3.  [2018-03-24T11:06:18,082][INFO ][logstash.pipeline        ] Starting pipeline {:pipeline_id=>"main", "pipeline.workers"=>4, "pipeline.batch.size"=>125, "pipeline.batch.delay"=>50}
4.  [2018-03-24T11:06:18,122][INFO ][logstash.pipeline        ] Pipeline started succesfully {:pipeline_id=>"main", :thread=>"#<Thread:0x6d4b91c4 sleep>"}
5.  [2018-03-24T11:06:18,133][INFO ][logstash.agent           ] Pipelines running {:count=>2, :pipelines=>[".monitoring-logstash", "main"]}

`AI写代码

因此,重新加载 Pipeline(数据处理管道)只需要不到 1 秒钟。这无疑是一个巨大的改进!

发送第一条测试数据

让我们再次使用之前的示例。

实际上,我创建了一个名为 samples.sh 的脚本,以便后续能够方便地添加更多测试场景:

vbnet 复制代码
`

1.  curl -XPOST "localhost:8080" -H "Content-Type: application/json" -d '{
2.    "test_case": "Address with text",
3.    "name": "Joe Smith",
4.    "address": {
5.      "number": "23",
6.      "street_name": "r verdiere",
7.      "city": "rochelle",
8.      "country": "France"
9.    }
10.  }'

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

运行该脚本后,得到的结果如下:

ini 复制代码
 `1.       "test_case" => "Address with text",
2.        "@version" => "1",
3.            "host" => "0:0:0:0:0:0:0:1",
4.      "@timestamp" => 2018-03-24T09:09:58.749Z,
5.            "name" => "Joe Smith",
6.         "address" => {
7.                 "city" => "rochelle",
8.          "street_name" => "r verdiere",
9.               "number" => "23",
10.              "country" => "France"
11.      },
12.         "headers" => {
13.           "content_length" => "158",
14.           "request_method" => "POST",
15.             "content_type" => "application/json",
16.                "http_host" => "localhost:8080",
17.             "request_path" => "/",
18.          "http_user_agent" => "curl/7.54.0",
19.              "request_uri" => "/",
20.             "http_version" => "HTTP/1.1",
21.              "http_accept" => "*/*"
22.      }
23.  }`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

使用地址信息查询 Elasticsearch

在本文开头,我们已经介绍了如何查询 Elasticsearch,从 BANO 数据集中获取有用的地址信息。

现在,让我们使用 elasticsearch-filter-plugin(Elasticsearch 过滤器插件),将这一查询能力集成到 Logstash 中:

ini 复制代码
`

1.  elasticsearch {
2.    query_template => "search-by-name.json"
3.    index => ".bano"
4.    fields => {
5.      "location" => "[location]"
6.      "address" => "[address]"
7.    }
8.    remove_field => ["headers", "host", "@version", "@timestamp"]
9.  }

`AI写代码

接下来,让我们解释几个之前尚未介绍过的新参数。

query_template 参数允许我们将 Elasticsearch 查询语句编写在一个外部文件中,而不必将整个查询压缩成一行并嵌入 Logstash 配置文件。

这样可以让代码更加清晰、易读,但也存在一个缺点。

当你修改 search-by-name.json 文件时,Logstash 不会检测到这一变更,因此也不会自动重新加载 Pipeline(数据处理管道)。

所以,我们需要采用一个小技巧:对 Pipeline 配置文件进行一次非常小的修改,以触发 Logstash 重新加载 Pipeline。

下面是 search-by-name.json 文件的内容:

css 复制代码
`

1.  {
2.    "size": 1,
3.    "query":{
4.      "bool": {
5.        "should": [
6.          {
7.            "match": {
8.              "address.number": "%{[address][number]}"
9.            }
10.          },
11.          {
12.            "match": {
13.              "address.street_name": "%{[address][street_name]}"
14.            }
15.          },
16.          {
17.            "match": {
18.              "address.city": "%{[address][city]}"
19.            }
20.          }
21.        ]
22.      }
23.    }
24.  }

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

看起来很熟悉,对吧?

index 参数用于指定我们希望查询的 Elasticsearch 索引或索引别名(Alias)。在这里,我们要查询的是 .bano 这个索引别名。

fields 参数用于指定我们希望从 Elasticsearch 查询结果中提取哪些字段,并将这些字段填充到当前的 Event(事件) 中(你也可以将其理解为文档)。

再次运行 samples.sh 脚本后,得到的结果如下:

ini 复制代码
`

1.  {
2.      "test_case" => "Address with text",
3.       "location" => {
4.          "lon" => -1.155167,
5.          "lat" => 46.157353
6.      },
7.           "name" => "Joe Smith",
8.        "address" => {
9.                 "city" => "La Rochelle",
10.          "street_name" => "Rue Verdière",
11.              "zipcode" => "17000",
12.               "number" => "23"
13.      }
14.  }

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

太棒了!成功了!

使用地理坐标查询 Elasticsearch

不过,我们还有另一个应用场景:根据地理位置搜索地址。

我们希望能够使用包含地理坐标信息的文档进行查询,例如下面这个示例。我已经将它添加到了 samples.sh 脚本中:

vbnet 复制代码
`

1.  curl -XPOST "localhost:8080" -H "Content-Type: application/json" -d '{
2.    "test_case": "Address with geo",
3.    "location": {
4.      "lat": 46.15735,
5.      "lon": -1.1551
6.    }
7.  }'

`AI写代码

运行后,得到的结果如下:

ini 复制代码
`

1.  {
2.      "test_case" => "Address with geo",
3.       "location" => {
4.          "lat" => 46.15735,
5.          "lon" => -1.1551
6.      }
7.  }

`AI写代码

我们需要在 Pipeline(数据处理管道)中添加一些条件判断逻辑,以便根据不同的输入数据,使用另一个 query_template(查询模板)来执行搜索:

ini 复制代码
 `1.    # We search by distance in that case
2.    elasticsearch {
3.      query_template => "search-by-geo.json"
4.      index => ".bano"
5.      fields => {
6.        "location" => "[location]"
7.        "address" => "[address]"
8.      }
9.      remove_field => ["headers", "host", "@version", "@timestamp"]
10.    }
11.  } else {
12.    # We search by address in that case
13.    elasticsearch {
14.      query_template => "search-by-name.json"
15.      index => ".bano"
16.      fields => {
17.        "location" => "[location]"
18.        "address" => "[address]"
19.      }
20.      remove_field => ["headers", "host", "@version", "@timestamp"]
21.    }
22.  }`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

search-by-geo.json 查询模板看起来也很熟悉:

css 复制代码
`

1.  {
2.    "size": 1,
3.    "query": {
4.      "bool": {
5.        "filter": {
6.          "geo_distance": {
7.            "distance": "1km",
8.            "location": {
9.              "lat": %{[location][lat]},
10.              "lon": %{[location][lon]}
11.            }
12.          }
13.        }
14.      }
15.    },
16.    "sort": [
17.      {
18.        "_geo_distance": {
19.          "location": {
20.            "lat": %{[location][lat]},
21.            "lon": %{[location][lon]}
22.          }
23.        }
24.      }
25.    ]
26.  }

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

现在,再次运行我们的示例脚本,得到的结果如下:

ini 复制代码
`

1.  {
2.      "test_case" => "Address with geo",
3.       "location" => {
4.          "lon" => -1.155167,
5.          "lat" => 46.157353
6.      },
7.        "address" => {
8.                 "city" => "La Rochelle",
9.          "street_name" => "Rue Verdière",
10.              "zipcode" => "17000",
11.               "number" => "23"
12.      }
13.  }

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

优化查询

我们之前也看到,每次查询都扫描整个数据集,效率可能非常低。

如果我们已经知道地址所属的省份编号,就可以通过直接查询对应的索引,而不是查询所有索引,来帮助 Elasticsearch 提高查询效率。

假设输入事件(Event)中包含邮政编码(Zipcode)或部分邮政编码。

例如,让我们在 samples.sh 脚本中添加以下几个测试场景:

vbnet 复制代码
 `1.    "test_case": "Address with geo and zipcode",
2.    "address": {
3.      "zipcode": "17000"
4.    },
5.    "location": {
6.      "lat": 46.15735,
7.      "lon": -1.1551
8.    }
9.  }'
10.  curl -XPOST "localhost:8080" -H "Content-Type: application/json" -d '{
11.    "test_case": "Address with geo and partial zipcode",
12.    "address": {
13.      "zipcode": "17"
14.    },
15.    "location": {
16.      "lat": 46.15735,
17.      "lon": -1.1551
18.    }
19.  }'`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

首先,在调用 Elasticsearch 之前,我们需要检查输入数据中是否包含邮政编码(zipcode)。

如果存在,就创建一个临时字段 dept,并通过截取操作只保留邮政编码的前两位数字。

在法国,这两位数字通常代表地址所属的省份编号:

ini 复制代码
`

1.  if [address][zipcode] {
2.    mutate { add_field => { "dept" => "%{[address][zipcode]}" } }
3.    truncate { 
4.      fields => ["dept"]
5.      length_bytes => 2
6.    }
7.  }

`AI写代码

接下来,让我们创建一个 index_suffix 字段:

ini 复制代码
`mutate { add_field => { "index_suffix" => "-%{dept}" } }`AI写代码

但是,如果输入数据中没有 zipcode(邮政编码)字段,我们就需要设置一些默认值:

dart 复制代码
`

1.  else {
2.    mutate { add_field => { "dept" => "" } }
3.    mutate { add_field => { "index_suffix" => "" } }
4.  }

`AI写代码

到目前为止,一切都很顺利。

不过,等等!我之前说过,在法国,事情永远不会那么简单。没错......我们还有编号为三位数字的省份!o_O

幸运的是,这些省份的编号都以 97 开头。因此,我们还需要在处理逻辑中考虑这种特殊情况:

ini 复制代码
`

1.  if [address][zipcode] {
2.    mutate { add_field => { "dept" => "%{[address][zipcode]}" } }
3.    truncate { 
4.      fields => ["dept"]
5.      length_bytes => 2
6.    }
7.    if [dept] == "97" {
8.      mutate { replace => { "dept" => "%{[address][zipcode]}" } }
9.      truncate { 
10.        fields => ["dept"]
11.        length_bytes => 3
12.      }
13.    }
14.    mutate { add_field => { "index_suffix" => "-%{dept}" } }
15.  } else {
16.    mutate { add_field => { "dept" => "" } }
17.    mutate { add_field => { "index_suffix" => "" } }
18.  }

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

现在,我们可以通过以下配置来修改索引名称:

ini 复制代码
`index => ".bano%{index_suffix}"`AI写代码

同时,我们还需要删除之前创建的两个临时字段:

ini 复制代码
`emove_field => ["headers", "host", "@version", "@timestamp", "index_suffix", "dept"]`AI写代码

最后,我们需要检查是否存在回归问题(Regression),尤其要重点验证省份编号为 974 的情况,确保修改后的逻辑仍然能够正确处理三位数的省份编号。

vbnet 复制代码
`

1.  curl -XPOST "localhost:8080" -H "Content-Type: application/json" -d '{
2.    "test_case": "Address with geo and zipcode in 974",
3.    "address": {
4.      "zipcode": "97400"
5.    },
6.    "location": {
7.      "lat": -21.214204,
8.      "lon": 55.361034
9.    }
10.  }'

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

结果确实返回了我们预期的正确值:

ini 复制代码
`

1.  {
2.      "test_case" => "Address with geo and zipcode in 974",
3.       "location" => {
4.          "lon" => 55.361034,
5.          "lat" => -21.214204
6.      },
7.        "address" => {
8.                 "city" => "Les Avirons",
9.          "street_name" => "Chemin des Acacias",
10.              "zipcode" => "97425",
11.               "number" => "13"
12.      }
13.  }

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)

下一步

请阅读下一篇文章,了解如何利用这一技术对现有数据进行增强和补全。

完整的 Logstash Pipeline(数据处理管道)

下面是我最终完成的完整 Logstash Pipeline 配置:

ini 复制代码
`

1.  input {
2.    http { }
3.  }

5.  filter {
6.    if [address][zipcode] {
7.      mutate { add_field => { "dept" => "%{[address][zipcode]}" } }
8.      truncate { 
9.        fields => ["dept"]
10.        length_bytes => 2
11.      }
12.      if [dept] == "97" {
13.        mutate { replace => { "dept" => "%{[address][zipcode]}" } }
14.        truncate { 
15.          fields => ["dept"]
16.          length_bytes => 3
17.        }
18.      }
19.      mutate { add_field => { "index_suffix" => "-%{dept}" } }
20.    } else {
21.      mutate { add_field => { "dept" => "" } }
22.      mutate { add_field => { "index_suffix" => "" } }
23.    }

25.    if [location][lat] and [location][lon] {
26.      elasticsearch {
27.        query_template => "search-by-geo.json"
28.        index => ".bano"
29.        fields => {
30.          "location" => "[location]"
31.          "address" => "[address]"
32.        }
33.        remove_field => ["headers", "host", "@version", "@timestamp", "index_suffix", "dept"]
34.      }
35.    } else {
36.      elasticsearch {
37.        query_template => "search-by-name.json"
38.        index => ".bano"
39.        fields => {
40.          "location" => "[location]"
41.          "address" => "[address]"
42.        }
43.        remove_field => ["headers", "host", "@version", "@timestamp", "index_suffix", "dept"]
44.      }
45.    }
46.  }

48.  output {
49.    stdout { codec => rubydebug }
50.  }

`AI写代码![](https://csdnimg.cn/release/blogv2/dist/pc/img/runCode/icon-arrowwhite.png)
相关推荐
可乐ea1 小时前
Git 基础设施重建:智能体规模开发下的读写解耦
大数据·git·elasticsearch·分布式存储·git基础设施·智能体规模开发·读写解耦
Elasticsearch3 小时前
使用 Elastic Stack 丰富邮政地址信息 —— 第一部分
elasticsearch
ofoxcoding1 天前
借助 CLAUDE.md 约束 Sonnet 5.5 多文件重构行为的提示词实践
大数据·elasticsearch·ai·重构
Elasticsearch1 天前
隆重推出 AlertZero:让你的告警队列实现 “收件箱清零” 式管理
elasticsearch
ly76891 天前
从 TF-IDF 到 BM25:Elasticsearch 相关性评分的字段长度归一化与 boost 调参边界
java·elasticsearch·query dsl·bm25·相关性评分
SL-staff1 天前
技术实践:将Excel自动升级为可搜索的知识节点(JVS平台实现)
mongodb·elasticsearch·知识图谱·jvs平台·语义解析·企业文档管理·excel结构化
Elasticsearch1 天前
2026 年唯一获得端点预防与响应(EPR)满分的厂商是 Elastic
elasticsearch
vx_Biye_Design1 天前
springboot旅游管理系统18006-计算机课程设计、毕业设计
java·spring boot·后端·python·elasticsearch·django·课程设计
ly76892 天前
倒排索引与 FST 在 Lucene 中的内存布局:segment、docValues 与词典压缩的工程取舍
elasticsearch·lucene·倒排索引·doc values·fst