【CE314】Computer Science NLP

Deadline: Please follow deadline on FASER

Build a text classifier on the IMDB sentiment classification dataset, you can use any classification method, but you must training your model on the first 40000 instances and testing your model on the last 10000 instances. The IMDB dataset will be uploaded on the moodle page for you to download.

Your code should include:

1: Read the file, incorporate the instances into the training set and testing set.

2: Pre-processing the text, you can choose whether you need stemming, removing stop words, removing non-alphabetical words. (Not all classification models need this step, it is OK if you think your model can perform better without this step, and you can give some justification in the report.)

3: Analysing the feature of the training set, report the linguistic features of the training dataset.

4: Build a text classification model, train your model on the training set and test your model on the test set.

5: Summarize the performance of your model (You can gain additional marks if you have some graph visualization).

6: (Optional) You can speculate how you can improve your works based on your proposed model.

After you build such a model and test on the test set, you should write a report (no longer than three pages in A4, with Arial 11 fonts) to summarize your work.

(You can use the existing algorithms on github or kaggle, but you must not directly copy and paste their code!

However, you are not allowed to use the Naïve Bayes algorithm and VADER classifier, which practiced in Lab 4)

Suggestion: some bonus points:

Have necessary comments on your code

Have proper reference on your report

Have graph visualization on your report

Investigate more evaluation methods, like not only show the P R F score, but also run multiple times and show the standard derivation on P R F (I am sure you can find more evaluation methods.)

Write your report like a mini-conference paper (you can learn from this paper:

  • Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy. 2016. Hierarchical Attention Networks for Document Classification. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages 1480--1489, San Diego, California. Association for Computational Linguistics.
相关推荐
惊讶的猫10 分钟前
《动手学大模型智能体》(Hands-on AI Agent)
人工智能
找方案13 分钟前
北京砸1亿支持智能体:Agent创业迎来黄金窗口期
大数据·人工智能·microsoft
Revolution6122 分钟前
多个 Agent 同时工作时,主 Agent 怎样接收队友结果
人工智能·llm·claude
不加辣椒29 分钟前
第8章:工具调用与多模态上下文集成
人工智能
葡萄城技术团队30 分钟前
三大适配场景:释放AI Coding真实落地价值(四)
人工智能
MomentYY36 分钟前
RAG 混合检索:关键词 + 语义
人工智能·agent·ai编程
Black蜡笔小新37 分钟前
EasyAIS+国标GB28181视频监控平台EasyCVR强强联动,全域视频AI识别能力落地!
大数据·人工智能·音视频
Microvision维视智造38 分钟前
智能工厂等级自检表:梯度培育,你的工厂在第几级?
人工智能·计算机视觉·视觉检测·机器视觉
时光不负努力43 分钟前
skill 定义 + 多个skill 协作
人工智能·openai
武子康1 小时前
MCP 2026-07-28 无状态核心之后:身份、任务、幂等与审计状态到底放在哪里?
人工智能·llm·mcp