nlp--最大匹配分词(计算召回率)

最大匹配算法是一种常见的中文分词算法,其核心思想是从左向右取词,以词典中最长的词为优先匹配。这里我将为你展示一个简单的最大匹配分词算法的实现,并结合输入任意句子、显示分词结果以及计算分词召回率。

代码 :

复制代码
# happy coding
# -*- coding: UTF-8 -*-
'''
@project:NLP
@auth:y1441206
@file:最大匹配法分词.py
@date:2024-06-30 16:08
'''
class MaxMatchSegmenter:
    def __init__(self, dictionary):
        self.dictionary = dictionary
        self.max_length = max(len(word) for word in dictionary)

    def segment(self, text):
        result = []
        index = 0
        n = len(text)

        while index < n:
            matched = False
            for length in range(self.max_length, 0, -1):
                if index + length <= n:
                    word = text[index:index+length]
                    if word in self.dictionary:
                        result.append(word)
                        index += length
                        matched = True
                        break
            if not matched:
                result.append(text[index])
                index += 1

        return result

def calculate_recall(reference, segmented):
    total_words = len(reference)
    correctly_segmented = sum(1 for word in segmented if word in reference)
    recall = correctly_segmented / total_words if total_words > 0 else 0
    return recall

# Example usage
if __name__ == "__main__":
    # Example dictionary
    dictionary = {"北京", "天安门", "广场", "国家", "博物馆", "人民", "大会堂", "长城"}

    # Example text to segment
    text = "北京天安门广场是中国的象征,国家博物馆和人民大会堂也在附近。"

    # Initialize segmenter with dictionary
    segmenter = MaxMatchSegmenter(dictionary)

    # Segment the text
    segmented_text = segmenter.segment(text)

    # Print segmented result
    print("分词结果:", " / ".join(segmented_text))

    # Example for calculating recall
    reference_segmentation = ["北京", "天安门广场", "是", "中国", "的", "象征", ",", "国家", "博物馆", "和", "人民大会堂", "也", "在", "附近", "。"]
    recall = calculate_recall(reference_segmentation, segmented_text)
    print("分词召回率:", recall)

运行结果 :

相关推荐
梦帮科技2 分钟前
可证伪性工程测评与生产部署红蓝对抗:压力测试、长尾鲁棒性与全生命周期安全质检守卫
人工智能·深度学习·神经网络·安全·机器学习·自然语言处理·压力测试
蜗牛互联网3 分钟前
Gemini 4 Argon的1M输出窗口与长程Agent工程边界
java·人工智能·后端
dtsola4 分钟前
一个人怎么指挥一支 AI 队伍
人工智能·程序员·ai创业·独立开发者·openclaw·一人公司·小遥claw
坤盾科技7 分钟前
用坤擎智能体搭一套企业级 AI 中台
人工智能·大模型·企业数字化·ai智能体·坤擎智能体
johnsong15 分钟前
幽灵协议:当AI推荐死去的SaaS,编码Agent控制层正在形成
大数据·人工智能
这张生成的图像能检测吗17 分钟前
(论文速读)DefectDiffu:基于一致性建模的少样本工业缺陷图像生成
人工智能·深度学习·计算机视觉·异常检测·少样本学习·扩散生成
7yewh27 分钟前
SLAM 从视觉里程计到建图(6)
数据结构·人工智能·机器人·自动驾驶·嵌入式
勤劳X码农28 分钟前
2026年AI配音做职场视频怎么选?
人工智能·音视频
YOLO数据集集合30 分钟前
玉米雄穗目标检测数据集 | 玉米雄穗 作物表型 智慧农业 无人机巡检 小样本检测9142期
人工智能·目标检测·无人机·玉米·玉米雄蕊·玉米雄穗
weixin_4045512437 分钟前
AI Native 架构建议:从零开始以 AI 为核心构建系统
人工智能·架构