【CE314】Computer Science NLP

Deadline: Please follow deadline on FASER

Build a text classifier on the IMDB sentiment classification dataset, you can use any classification method, but you must training your model on the first 40000 instances and testing your model on the last 10000 instances. The IMDB dataset will be uploaded on the moodle page for you to download.

Your code should include:

1: Read the file, incorporate the instances into the training set and testing set.

2: Pre-processing the text, you can choose whether you need stemming, removing stop words, removing non-alphabetical words. (Not all classification models need this step, it is OK if you think your model can perform better without this step, and you can give some justification in the report.)

3: Analysing the feature of the training set, report the linguistic features of the training dataset.

4: Build a text classification model, train your model on the training set and test your model on the test set.

5: Summarize the performance of your model (You can gain additional marks if you have some graph visualization).

6: (Optional) You can speculate how you can improve your works based on your proposed model.

After you build such a model and test on the test set, you should write a report (no longer than three pages in A4, with Arial 11 fonts) to summarize your work.

(You can use the existing algorithms on github or kaggle, but you must not directly copy and paste their code!

However, you are not allowed to use the Naïve Bayes algorithm and VADER classifier, which practiced in Lab 4)

Suggestion: some bonus points:

Have necessary comments on your code

Have proper reference on your report

Have graph visualization on your report

Investigate more evaluation methods, like not only show the P R F score, but also run multiple times and show the standard derivation on P R F (I am sure you can find more evaluation methods.)

Write your report like a mini-conference paper (you can learn from this paper:

  • Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy. 2016. Hierarchical Attention Networks for Document Classification. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages 1480--1489, San Diego, California. Association for Computational Linguistics.
相关推荐
想会飞的蒲公英2 分钟前
用 Streamlit 给文本分类模型做一个演示页面
人工智能·python·机器学习
老徐聊GEO6 分钟前
亲测有效的AI品牌检测公司案例分享
大数据·人工智能·python
IT智慧客07317 分钟前
Vibe Coding 时代:Vue 消失了还是 React 太强?
人工智能
韭菜学长13 分钟前
科技型中小企业如何申请政府补贴?申报指南
大数据·人工智能
触底反弹16 分钟前
🔥 AI 写代码总翻车?这套实战方法论救了我
人工智能·面试·程序员
风栖柳白杨17 分钟前
【面试】AI算法工程师_空白自测版本
人工智能·算法·面试
默大老板是在下19 分钟前
信息深度加工框架:如何把“看过”变成“能调用”
人工智能
隔窗听雨眠21 分钟前
AI Agent可观测性:破解多步推理黑盒
人工智能
延凡科技23 分钟前
智慧燃气解决方案:管道煤气监管应用管理系统(三维GIS+IoT+大数据落地实战)
大数据·人工智能·科技·物联网·安全
水如烟34 分钟前
孤能子视角:梁文锋的场,一个量化交易者的本能
人工智能