View Proposal


Proposer
Beatrice Alex
Title
Advancing Language Modelling for Scottish Gaelic, a Low-resource Language
Goal
This project aims to move Scottish Gaelic language modelling beyond evaluation toward active improvement. Building on GaelEval, the first multi-task Gaelic benchmark, it will investigate methods for enhancing large language model (LLM) capability on Gaelic, such as targeted fine-tuning or post-training on existing digitised data, with data augmentation as one avenue for addressing resource scarcity. While the ultimate goal is to improve LLM performance on Gaelic tasks using GaelEval, the project's value lies as much in understanding which methods help and why. A well-executed study that explores the limits of current approaches is a strong outcome in its own right.
Description
Scottish Gaelic is a morpho-syntactically rich minority language with a small number of active speakers, limited digital resources and no dedicated large language model of its own. While multilingual LLMs exhibit some capability in languages they were not explicitly trained on, GaelEval [1] — the first multi-task benchmark for Gaelic — has shown that performance remains uneven across linguistic, cultural and translation tasks, with proprietary frontier models consistently outperforming open-weight ones. Drawing on established approaches to modelling other low-resource languages (e.g. Basque and Irish [2–4]), this project will build on GaelEval to investigate methods for enhancing LLM performance on Gaelic, working from existing digitised Gaelic datasets. The work directly confronts the core challenge of building capability for a morpho-syntactically rich language with few speakers and scarce digital resources. Students should contact the supervisor if they are interested in pursuing research in this area, ideally with a more concrete idea of which aspect of the project they would like to focus on. Students may also propose their own project in the area of low-resource languages and NLP/AI, in which case they should contact the supervisor with details of their project idea and the dataset they would like to work with.
Resources
[1] Devine, Peter, et al. "GaelEval: Benchmarking LLM Performance for Scottish Gaelic.” LLMs4SSH, LREC 2026 (2026). [2] Etxaniz, Julen, et al. "Latxa: An open language model and evaluation suite for Basque." Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. [3] https://www.latamgpt.org/en [4] Barry, James, et al. "gaBERT—an Irish language model." Proceedings of the Thirteenth Language Resources and Evaluation Conference. 2022.
Background
Url
External Link
Difficulty Level
Moderate
Ethical Approval
None
Number Of Students
2
Supervisor
Beatrice Alex
Keywords
natural language processing, nlp, machine learning, large language models, evaluation, low-resource language
Degrees
Bachelor of Science in Computer Systems
Master of Science in Artificial Intelligence
Master of Science in Artificial Intelligence with SMI
Master of Science in Data Science
Bachelor of Science in Computing Science
BSc Data Sciences