View Proposal
-
Proposer
-
John See
-
Title
-
MY-Guard: Localising AI Guardrails for Malaysian Socio-cultural Contexts (TAKEN)
-
Goal
-
The aim of this project is to investigate whether existing guardrails actually work in the Malaysian context where socio-cultural nuances may be very different from what they were originally developed. To achieve this, a Malaysian-centric safety benchmark should be built whilst attempting to adapt existing safety models through model fine-tuning to better recognise these contexts. Some form of validation (qualitative and/or quantitative) is required.
-
Description
- This project will evaluate an open-source guardrail architecture based on open source libraries and models (e.g. Nvidia NeMo Guardrails, and Llama Guard 3 model) against a purpose-built Malaysia safety benchmark covering multilingual, cultural and socially senstive scenarios. This entails the creation of a Malaysian-specific safety dataset (can be synthetically created), which is then used to adapt the safety model through supervised fine-tuning (SFT) with LoRA (or its variants). For validation, the adapted guardrail will be evaluated against the original system to quantify improvements in safety and contextual appropriateness while maintaining responses to legitimate queries.
*This project requires a deep understanding of socio-cultural issues, taboos and societal norms in Malaysia.
- Resources
-
High-performance compute for model training
-
Background
-
In the present age where artificial intelligence (AI) has permeated various aspects of our lives, spearheaded by advances in Large Language Models (LLM), there is an increasing concern towards making LLMs safe. Existing "guardrails" (or safeguards designed to keep AI systems operating safely, responsibly, and within defined behavioral boundaries) may not fully capture the socio-cultural nuances of the Malaysian society. For instance, users may decide to ask AI to generate offensive statements that may be used maliciously. In a different setting, an LLM response may misunderstood questions that contain complex code-switching (a mix of different languages in a statement) especially in a multi-cultural society like Malaysia. In a nutshell, these guardrails developed primarily from English-language and Western-centric safety datasets may not adequately capture the linguistic nuances, cultural sensitivities, social norms and other context-specific issues encountered by Malaysian users.
Malaysia is also developing its AI ecosystem around responsible AI. In particular, the National AI Office identifies responsible, safe, secure, ethical and trustworthy AI as key governance objectives.
-
Url
-
-
Difficulty Level
-
High
-
Ethical Approval
-
None
-
Number Of Students
-
1
-
Supervisor
-
John See
-
Keywords
-
responsible ai, llm, guardrails, safety models
-
Degrees
-
Bachelor of Science in Computing Science