【CE314】Computer Science NLP

Deadline: Please follow deadline on FASER

Build a text classifier on the IMDB sentiment classification dataset, you can use any classification method, but you must training your model on the first 40000 instances and testing your model on the last 10000 instances. The IMDB dataset will be uploaded on the moodle page for you to download.

Your code should include:

1: Read the file, incorporate the instances into the training set and testing set.

2: Pre-processing the text, you can choose whether you need stemming, removing stop words, removing non-alphabetical words. (Not all classification models need this step, it is OK if you think your model can perform better without this step, and you can give some justification in the report.)

3: Analysing the feature of the training set, report the linguistic features of the training dataset.

4: Build a text classification model, train your model on the training set and test your model on the test set.

5: Summarize the performance of your model (You can gain additional marks if you have some graph visualization).

6: (Optional) You can speculate how you can improve your works based on your proposed model.

After you build such a model and test on the test set, you should write a report (no longer than three pages in A4, with Arial 11 fonts) to summarize your work.

(You can use the existing algorithms on github or kaggle, but you must not directly copy and paste their code!

However, you are not allowed to use the Naïve Bayes algorithm and VADER classifier, which practiced in Lab 4)

Suggestion: some bonus points:

Have necessary comments on your code

Have proper reference on your report

Have graph visualization on your report

Investigate more evaluation methods, like not only show the P R F score, but also run multiple times and show the standard derivation on P R F (I am sure you can find more evaluation methods.)

Write your report like a mini-conference paper (you can learn from this paper:

  • Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy. 2016. Hierarchical Attention Networks for Document Classification. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages 1480--1489, San Diego, California. Association for Computational Linguistics.
相关推荐
Light Gao8 小时前
RNN 原理、架构与一次完整手算
人工智能·rnn·深度学习·神经网络·embedding
weixin_438338518 小时前
CUDA 安装理解
人工智能·深度学习
喵本喵叁肆8 小时前
diffprivlib 实战(四):训练差分隐私机器学习模型——sklearn 一行换 import 就是 DP 版
人工智能·机器学习·sklearn
郝学胜-神的一滴8 小时前
AI 编程智能体 02:AI智能体到底是什么
开发语言·人工智能·python·程序人生·pycharm
天一生水water8 小时前
超参分析入门教程
人工智能
LaughingZhu8 小时前
Product Hunt 每日热榜 | 2026-10-01
数据库·人工智能·深度学习·神经网络·搜索引擎
橘和柠8 小时前
CUDA 与 N 卡驱动安装
人工智能
悟天特斯8 小时前
边缘计算:让智慧园区的治理能力“下沉“到最后一公里
人工智能·边缘计算
宋哥转AI8 小时前
AgentScope Java 实战 05:Spring Boot 整合——11 个官方 starter 的自动装配路线
人工智能·ai·ai编程
I Am a robert girl8 小时前
从一张图到碎裂瞬间:FracGen 如何用物理信号“导演“物体撕裂
人工智能·深度学习·计算机视觉·生成模型·视频生成·物理仿真·断裂模拟