03停用词过滤

1.任务目标

本任务承接 Task02 文本词频统计项目,属于NLP自然语言预处理基础核心任务 。在原有文本读取、清洗、词频统计的基础上,新增停用词过滤功能,剔除无实际语义的高频虚词,筛选出有实际意义的核心高频词汇,实现英文文本精细化词频分析。

2.步骤

步骤1:新增核心功能函数

  1. filter_stopwords():接收单词列表与停用词集合,过滤所有停用词,输出纯实词列表;
python 复制代码
# 停用词过滤
def filter_stopwords(words, stop_set):
    # 定义空列表,存储过滤后的有效单词
    filtered = []
    # 遍历所有单词,只保留非停用词
    for word in words:
        if word not in stop_set:
            filtered.append(word)
    return filtered
  1. count_words():遍历过滤后的单词,统计词频生成精准词频字典;
python 复制代码
# 词频统计函数
def count_words(words):
    counts = {}
    for w in words:
        # 单词存在则+1,不存在则初始化为0再+1
        counts[w] = counts.get(w,0)+1
    return counts
  1. get_top_n():对词频字典降序排序,提取指定数量的高频核心词汇。
python 复制代码
# 提取TopN高频词并降序排序
def get_top_n(count_dict, n):
    # 按词频数值降序排序
    sorted_items = sorted(count_dict.items(), key=lambda x:x[1], reverse=True)
    # 截取前n个高频词
    return sorted_items[:n]

步骤2:复用

python 复制代码
# 读取本地文本文件函数
def read_file(filename):
    # 以utf-8编码打开文件,避免乱码,with自动关闭文件,更安全
    with open(filename, encoding = "utf-8") as f:
        return f.read()

# 文本清洗、转小写、去标点、切割单词
def clean_split(text):
    # 归一化
    text_lower = text.lower()
    # 去标点
    text_clean = ""
    for ch in text_lower:
        if ch not in string.punctuation:
            text_clean += ch
    # 切割
    words = text_clean.split()
    return words

步骤3:编写主程序流水线

python 复制代码
if __name__ == "__main__":
    # 定义英文通用停用词集合(set查询效率O(1))
    stop_words = {"the","a","an","is","are","of","in","on","at","to","and","or","but","as","that","this"}
    # 1.读取文本
    content = read_file("article.txt")
    # 2.清洗文本、切割单词
    word_list = clean_split(content)
    # 3.过滤停用词,得到有效实词列表
    filtered_words = filter_stopwords(word_list, stop_words)
    # 4.统计精准词频
    word_counts = count_words(filtered_words)
    # 5.获取Top10高频核心词汇
    top_10 = get_top_n(word_counts,10)
    # 6.打印结果
    print("过滤停用词后的Top10高频实词:")
    for word, cnt in top_10:
        print(f"{word}: {cnt}次")

3.总结

1、停用词集合为什么用 set 而非 list?

list 查询时间复杂度 O(n),遍历效率低;set 基于哈希表,成员查询 O(1),速度极快,适合大规模文本过滤。

2、列表和字典的核心区别?

列表是有序容器,通过下标取值,支持排序切片;字典是键值对容器,通过 key 取值,无序、不能排序、不能切片。

3、items() 的作用是什么?

将字典中所有键值对提取出来,转换为可排序、可遍历的列表结构,解决字典无法排序的问题。

相关推荐
言乐62 小时前
HTML视频审核模型
python·django·virtualenv·pygame·tornado
言乐63 小时前
Python根据无法识别搜索词找出可能输入内容模型
开发语言·python·django·virtualenv·pygame
by209994 小时前
学会使用std::string类,并理解其内部是如何管理字符串的详细阐述(上)
c++·笔记·字符串·类和对象·string
yl45304 小时前
硫酸泄露处理生产商怎么选才够专业
大数据·人工智能·python
笨笨饿4 小时前
140_AI新手村MCP与Skills是干嘛的
开发语言·人工智能·python·stm32·单片机·嵌入式硬件·物联网
笑鸿的学习笔记5 小时前
C++笔记之大块顺序写
java·c++·笔记
for_ever_love__5 小时前
机器学习入门——手写线性回归与梯度下降
人工智能·python·学习·机器学习·线性回归
打工仔折腾 AI5 小时前
从 Demo 到生产级 Agent:8 个关键设计机制与 Python 实现拆解
java·jvm·人工智能·后端·python·langchain·ai agent 实战
I Am a robert girl5 小时前
当传感器学会“说谎“:拆解可靠性门控的稀疏惯性动捕融合
python·姿态估计·传感器融合·惯性动捕·imu传感器·可靠性门控·可穿戴计算
李航19835 小时前
AI定制柜建模,需要详细的建模规范和标准流程
人工智能·python·计算机视觉·ai·ai编程