【CE314】Computer Science NLP

Deadline: Please follow deadline on FASER

Build a text classifier on the IMDB sentiment classification dataset, you can use any classification method, but you must training your model on the first 40000 instances and testing your model on the last 10000 instances. The IMDB dataset will be uploaded on the moodle page for you to download.

Your code should include:

1: Read the file, incorporate the instances into the training set and testing set.

2: Pre-processing the text, you can choose whether you need stemming, removing stop words, removing non-alphabetical words. (Not all classification models need this step, it is OK if you think your model can perform better without this step, and you can give some justification in the report.)

3: Analysing the feature of the training set, report the linguistic features of the training dataset.

4: Build a text classification model, train your model on the training set and test your model on the test set.

5: Summarize the performance of your model (You can gain additional marks if you have some graph visualization).

6: (Optional) You can speculate how you can improve your works based on your proposed model.

After you build such a model and test on the test set, you should write a report (no longer than three pages in A4, with Arial 11 fonts) to summarize your work.

(You can use the existing algorithms on github or kaggle, but you must not directly copy and paste their code!

However, you are not allowed to use the Naïve Bayes algorithm and VADER classifier, which practiced in Lab 4)

Suggestion: some bonus points:

Have necessary comments on your code

Have proper reference on your report

Have graph visualization on your report

Investigate more evaluation methods, like not only show the P R F score, but also run multiple times and show the standard derivation on P R F (I am sure you can find more evaluation methods.)

Write your report like a mini-conference paper (you can learn from this paper:

  • Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy. 2016. Hierarchical Attention Networks for Document Classification. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages 1480--1489, San Diego, California. Association for Computational Linguistics.
相关推荐
童园管理札记几秒前
CSDN 学前入门高质量指南:从零搭建编程学习体系
人工智能·经验分享·职场和发展·生活·学习方法
ShallWeL1 分钟前
【Agent工程】(15)—— 评测集与回归门禁
人工智能·agent·工作流
罗西的思考3 分钟前
[Agent Memory / 强化学习] MemPO源码学习笔记 —(1)— 总体
人工智能·笔记·深度学习·学习·机器学习
量化吞吐机12 分钟前
会写代码学量化,先把数据、规则和执行连清楚
人工智能·python
天远API13 分钟前
零信任架构实战:基于天远手机空号检测V即时版构建自动化新客入驻网关
人工智能·智能手机·架构·自动化
海盗123419 分钟前
AI 新闻日报 2026-09-10:DeepSeek V4.1 Flash 上线降价60%、京东狼族机器人军团、优必选联姻沐曦造芯
人工智能·机器人
词却26 分钟前
深度学习入门:卷积神经网络与 MNIST 手写数字识别
人工智能·深度学习·cnn
明志数科27 分钟前
流水线数据采集工程实践:MES产线监控数据与机器人Ego训练数据的本质区别与选型指南
人工智能
DogDaoDao30 分钟前
DexPIE:让人类手把手教灵巧手“回炉重造“——真实世界后训练 RL 深度拆解
深度学习·机器学习·机器人·人形机器人·运动轨迹·模仿学习·dexpie
乐迪信息31 分钟前
航道船舶逆行AI识别,港口安全智能告警系统
大数据·前端·人工智能·安全·计算机视觉·音视频