【CE314】Computer Science NLP

Deadline: Please follow deadline on FASER

Build a text classifier on the IMDB sentiment classification dataset, you can use any classification method, but you must training your model on the first 40000 instances and testing your model on the last 10000 instances. The IMDB dataset will be uploaded on the moodle page for you to download.

Your code should include:

1: Read the file, incorporate the instances into the training set and testing set.

2: Pre-processing the text, you can choose whether you need stemming, removing stop words, removing non-alphabetical words. (Not all classification models need this step, it is OK if you think your model can perform better without this step, and you can give some justification in the report.)

3: Analysing the feature of the training set, report the linguistic features of the training dataset.

4: Build a text classification model, train your model on the training set and test your model on the test set.

5: Summarize the performance of your model (You can gain additional marks if you have some graph visualization).

6: (Optional) You can speculate how you can improve your works based on your proposed model.

After you build such a model and test on the test set, you should write a report (no longer than three pages in A4, with Arial 11 fonts) to summarize your work.

(You can use the existing algorithms on github or kaggle, but you must not directly copy and paste their code!

However, you are not allowed to use the Naïve Bayes algorithm and VADER classifier, which practiced in Lab 4)

Suggestion: some bonus points:

Have necessary comments on your code

Have proper reference on your report

Have graph visualization on your report

Investigate more evaluation methods, like not only show the P R F score, but also run multiple times and show the standard derivation on P R F (I am sure you can find more evaluation methods.)

Write your report like a mini-conference paper (you can learn from this paper:

  • Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy. 2016. Hierarchical Attention Networks for Document Classification. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages 1480--1489, San Diego, California. Association for Computational Linguistics.
相关推荐
zhangfeng11334 分钟前
ATK(华为算子测试平台)详细介绍 CANN(Compute Architecture for Neural Networks,神经网络计算架构
人工智能·华为·ai编程·npu·cann
吉安特尔雅5 分钟前
业务级招聘自动化落地实践:桌面端 AI 招聘组件的部署、集成与风险管控要点
人工智能·ai招聘哪家靠谱
老王以为6 分钟前
走进 AI Agent 第四篇(上):知识获取管道——RAG 基础
前端·人工智能·全栈
唐璜Taro6 分钟前
Agent Harness 系列 · 第 1 篇|Agent 不只是换一个更强的模型
人工智能·python
海宇大数据16 分钟前
零信任架构实战:基于海宇学历证书核验构建自动化教育资质网关
人工智能·ai·工具分享
xx_xxxxx_19 分钟前
论文阅读-Search-R1
人工智能·深度学习·机器学习·强化学习
代码方舟24 分钟前
零信任架构实战:基于天远二手车VIN估值构建自动化车险承保网关
人工智能·ai·工具分享
灵海之森30 分钟前
大模型时代对话意图识别全景实践:从80个细粒度小类到生产级架构演进
人工智能
jimmyleeee31 分钟前
大模型安全之二十二:AI 安全实践:真实案例研究与经验教训
人工智能·安全
Ivanqhz1 小时前
Slope One 算法详解:与矩阵分解、共现矩阵的异同
java·服务器·网络·人工智能·深度学习