01词频统计器

1.任务目标

读一篇考研英语真题 文章,忽略大小写和标点,统计每个单词出现次数,输出前10个高频词及次数。

2.步骤

  1. 准备语料 :找一篇考研英语阅读真题原文,存成 article.txt,与代码同目录。存文件时编码选 UTF-8

    复制代码
    task01_wordcount/
    ├─ wordcount.py      
    └─ article.txt      
  2. 读取文件:用 open() 读入全文;先 print 出来确认读到了原文,再往下走

    python 复制代码
    # 读取文章
    f = open("article.txt", encoding = "utf-8") # 读取文件
    text = f.read()         # 把 f 里的内容读出来,存进变量 text
    print(text)           # 把 text 打印到终端
  3. 清洗 + 切词:

    1. 转小写 lower()

      python 复制代码
      text_lower = text.lower() # 将所有单词转成小写,并存进text_lower
      print(text_lower) # 验证
    2. 去标点,利用string.punctuation

      python 复制代码
      text_clean = ""
      for ch in text_lower:
          if ch not in string.punctuation:
              text_clean += ch
      print (text_clean) # 验证
    3. split() 成单词列表

      python 复制代码
      # 第四步 将单词切分放入列表
      words = text_clean.split() # 切割成单词,返回列表words
      print(words) # 验证
  4. 统计词频:

    python 复制代码
    counts = {}
    for word in words:
        if word in counts:        # 这个单词已经见过了
            counts[word] += 1
        else:                     # 头一次见
            counts[word] = 1
    print(counts) #验证
  5. 排序取前10:

    python 复制代码
    top_10 = sorted(counts.items(), key = lambda item : item[1], reverse = True) [:10]
    print(top_10) # 验证
相关推荐
打工仔折腾 AI2 小时前
从Attention到BERT:双向预训练语言模型到底解决了什么问题
人工智能·后端·python·深度学习·语言模型·bert
迅猛龙办公室2 小时前
实现第一个python程序(HelloWorld)
开发语言·python
有范先生2 小时前
这次户外断粮危机后,我更在意主食能不能扛住极端环境了
笔记
拉格朗日(Lagrange)3 小时前
【第2 章】WorkBuddy 从入门到高手
开发语言·python
喜欢打篮球的普通人3 小时前
MiniMind 学习笔记(十二):Pretrain 实操——从版本梳理到 8GB 显卡上的真实训练
人工智能·笔记·学习
老歌老听老掉牙4 小时前
斜抛运动问题分析:给定最大高度与墙面位置的轨迹与时间求解
python·斜抛运动
言乐64 小时前
Python自动去除水印
开发语言·python·django·virtualenv·pygame
czq_26867194874 小时前
Python打卡第31天
开发语言·python
invicinble4 小时前
python 编程语言 认识维度
开发语言·数据库·python
AC赳赳老秦4 小时前
数据采集全链路审计留痕:用 OpenClaw 实现合规审计与追溯
开发语言·汇编·python·php·swift·deepseek·openclaw