【CE314】Computer Science NLP

Deadline: Please follow deadline on FASER

Build a text classifier on the IMDB sentiment classification dataset, you can use any classification method, but you must training your model on the first 40000 instances and testing your model on the last 10000 instances. The IMDB dataset will be uploaded on the moodle page for you to download.

Your code should include:

1: Read the file, incorporate the instances into the training set and testing set.

2: Pre-processing the text, you can choose whether you need stemming, removing stop words, removing non-alphabetical words. (Not all classification models need this step, it is OK if you think your model can perform better without this step, and you can give some justification in the report.)

3: Analysing the feature of the training set, report the linguistic features of the training dataset.

4: Build a text classification model, train your model on the training set and test your model on the test set.

5: Summarize the performance of your model (You can gain additional marks if you have some graph visualization).

6: (Optional) You can speculate how you can improve your works based on your proposed model.

After you build such a model and test on the test set, you should write a report (no longer than three pages in A4, with Arial 11 fonts) to summarize your work.

(You can use the existing algorithms on github or kaggle, but you must not directly copy and paste their code!

However, you are not allowed to use the Naïve Bayes algorithm and VADER classifier, which practiced in Lab 4)

Suggestion: some bonus points:

Have necessary comments on your code

Have proper reference on your report

Have graph visualization on your report

Investigate more evaluation methods, like not only show the P R F score, but also run multiple times and show the standard derivation on P R F (I am sure you can find more evaluation methods.)

Write your report like a mini-conference paper (you can learn from this paper:

  • Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy. 2016. Hierarchical Attention Networks for Document Classification. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages 1480--1489, San Diego, California. Association for Computational Linguistics.
相关推荐
2601_962177304 分钟前
小白安装Claude Code完整教程:Windows从零装好并接入Crazyrouter(附403解决方法)
人工智能·深度学习·目标检测·机器学习·数据挖掘
m4Rk_6 分钟前
【论文阅读】Agent 记忆机制(82):ReasoningBank——从成功与失败经验中沉淀可复用的推理记忆
论文阅读·人工智能·学习·开源·github
XMAIPC_Robot7 分钟前
为什么大型储能需要边缘 AI 协调控制器?RK3588+FPGA 高速采集方案
人工智能·嵌入式硬件·fpga开发·arm+fpga·rk3588+fpga·协调控制器
xx_xxxxx_8 分钟前
论文阅读-CoTTA
人工智能·深度学习·机器学习
RPAdaren15 分钟前
金融政企高合规场景,该选现场编排还是流程库调用型 AI Agent
人工智能
泥人张19 分钟前
踩坑无数换来的教训:指挥AI开发App,这几点你必须知道
人工智能
蓝速科技21 分钟前
政务自助终端信创选型与无人值守落地方案
android·大数据·数据库·人工智能·科技·技术分享·政务
“AI国潮设计-小江”23 分钟前
[AIGC实战] 基于Stable Diffusion的潮汕非遗IP自动化生成工作流(附Python批量处理脚本)
开发语言·人工智能·python·prompt·aigc
szxinmai主板定制专家26 分钟前
RK3576+CODESYS+RK182X+FPGA异构架构:半导体精密设备一体化控制器方案
人工智能·fpga开发·架构·rk3576+codesys
EatFan33 分钟前
GitHub Trending 三连观察:AI 编码智能体进入“周边生态”竞争阶段
人工智能·驱动开发·github·mcp·superpowers·github trending·spec kit