View Proposal


Proposer
Ian Tan
Title
Small Language Models for Malaysian Social Media Opinion Analysis
Goal
This is an exploratory project where the expectation is to be able to train and build a new model specifically for public opinion towards specific social media content.
Description
The project is to build a Small Language Models (SLMs) tailored for Malaysian social media opinion analysis where it should be able to analyse mixed-language, informal text (Bahasa Rojak, text speak, and/or Gen Z lingo). The intention is to use the NeoBERT, a newer improved version of BERT (https://arxiv.org/html/2502.19587v1) to build the model. Existingly, there are a few other models available, such as TinyLLama-Malay (https://arxiv.org/pdf/2410.06973), and Mistral-based Model (https://arxiv.org/html/2401.13565v2). These are supported by datasets such as Malaysia Tweets Sentiment Dataset (https://huggingface.co/datasets/kaiimran/malaysia-tweets-sentiment), and Annotated dataset for sentiment analysis and sarcasm detection: Bilingual code-mixed English-Malay social media data in the public security domain (https://www.sciencedirect.com/science/article/pii/S2352340924006309).
Resources
Minimally, an A5000 GPU (according to NeoBERT requirements)
Background
Url
Difficulty Level
High
Ethical Approval
None
Number Of Students
1
Supervisor
Ian Tan
Keywords
slm, neobert, social media analytics, opinion analysis
Degrees
Bachelor of Science in Computing Science