View Proposal


Proposer
Beatrice Alex
Title
Clinical Natural Language Processing: Extracting Phenotypic Knowledge from Clinical Text
Goal
These projects aim to develop and evaluate NLP methods for extracting clinical phenotypes from unstructured text, including radiology reports, discharge summaries, referral letters, ICU notes etc. with an emphasis on robustness, transferability across clinical contexts, and trustworthiness for real-world deployment. Students can select one of four topics: improving extraction of rare but clinically important phenotypes; improving phenotyping of MRI (versus CT) reports or reports for younger patients; handling hedging and uncertainty in clinical language; or detecting and mitigating hallucination in clinical NLP. The work undertaken in these projects draws on real clinical data (e.g. NHS Scotland brain radiology reports) or established benchmarks (MIMIC-IV, MedQA/MultiMedQA), and targets known challenges in the field.
Description
Clinical unstructured text, e.g. radiology reports, discharge summaries, referral letters or ICU notes, is a rich and underused source of phenotypic information in healthcare. Natural language processing (NLP) offers tools and techniques to unlock this information at scale, but significant challenges remain: rare phenotypes (in this context usually observed diseases or symptoms) are severely under-represented, models trained on one clinical context often fail in another, clinical text is hedged and uncertain and can produce unsupported or fabricated findings. Moreover, the lack of rigorous dataset documentation makes replication and deployment challenging [1, 2]. Projects in this area will develop and evaluate NLP methods for clinical phenotype extraction, with a focus on robustness, transferability and trustworthiness. The work will make use of real clinical datasets, e.g. NHS Scotland brain radiology reports [3, 4], established datasets/benchmarks such as MIMIC-IV [5, 6] or MedQA/MultiMedQA [7, 8]; other existing benchmarks may also be used. This listing covers four separate research strands and students should indicate which is of interest to them: • Improving the classification of rare but clinically important phenotypes • Improving phenotype classification of MRI scan reports (versus CT) or that of reports for younger patients • Modelling uncertainty in text (e.g. suspected versus confirmed findings) • Hallucination detection and mitigation in clinical NLP Some topics may require additional annotation of gold-standard data or identifying a suitable existing annotated dataset. Students taking on one of these projects need to have training in NLP (namely F20NL/F21NL) with experience in working with textual data. An interest in working with and processing clinical or biomedical text is desirable. Students will need to undergo relevant training on GDPR and working safely with de-identified patient data as part of their project to obtain access to datasets. Students who would like to develop their own proposal in the area of clinical NLP can also be considered but should contact the supervisor with more details on their project idea and the dataset they would like to work with.
Resources
[1] Casey, Arlene, et al. "A systematic review of natural language processing applied to radiology reports." BMC medical informatics and decision making 21.1 (2021): 179. [2] Rahman, Fahrurrozi, et al. "Natural language processing for geriatric syndromes: a systematic review of methods, applications, and challenges." BMC Medical Informatics and Decision Making (2026). [3] Alex, Beatrice, et al. "GS-BrainText: A Multi-Site Brain Imaging Report Dataset from Generation Scotland for Clinical Natural Language Processing Development and Validation.” Clinical NLP Workshop, LREC 2026 (2026). [4] Camilleri, Michael PJ et al. "A large dataset of brain imaging linked to health systems data: curation and access to a whole system national cohort from NHS Scotland." GigaScience (2026): giag072. [5] Johnson, Alistair, Lucas Bulgarelli, Tom Pollard, Steven Horng, Leo Anthony Celi, and Roger Mark. “MIMIC-IV." PhysioNet. Available online at: https://physionet. org/content/mimiciv/1.0/(accessed August 23, 2021) (2020). [6] Du, Hongbo et al., Possible or Definite? A Benchmark for Evaluating Diagnostic Uncertainty Preservation in Clinical Text, arXiv: https://arxiv.org/pdf/2606.18471, (2026). [7] Jin, Di, et al. "What disease does this patient have? A large-scale open domain question answering dataset from medical exams." Applied Sciences 11.14 (2021): 6421. [8] Singhal, Karan, et al. "Large language models encode clinical knowledge." Nature 620.7972 (2023): 172–180.
Background
Url
Difficulty Level
Moderate
Ethical Approval
Full
Number Of Students
4
Supervisor
Beatrice Alex
Keywords
natural language processing, nlp, clinical text, machine learning, large language models, evaluation
Degrees
Bachelor of Science in Computer Science
Master of Science in Artificial Intelligence
Master of Science in Artificial Intelligence with SMI
Master of Science in Data Science
Bachelor of Science in Computing Science
BSc Data Sciences