View Proposal


Proposer
Ian Tan
Title
Human, Too Human: Investigating Sequential Priming and Anchoring Biases in Large Language Model Based Automated Essay Grading
Goal
To evaluate whether LLMs exhibit human-like sequential anchoring and contrast biases when evaluating student essays in series.
Description
Building on earlier findings (previous year's FYP) regarding prompt engineering and human-LLM agreement in Automated Essay Grading (AEG), this project investigates the impact of sequence dynamics and context priming on LLM scoring behaviour. While automated scoring systems aim to eliminate human limitations such as fatigue and subjectiveness, generative models evaluated on sequential inputs may replicate human biases. Specifically, human evaluators frequently suffer from anchoring bias, where an exceptionally strong or weak initial submission anchors the benchmark for subsequent essays, where identical essays receive systematically lower scores following high-quality submissions than following low-quality ones. This study establishes an experimental framework to test whether LLMs (will determine the various LLMs to study during the project) exhibit sequence dependent variance during essay grading. The proposed methodology involves presenting identical benchmark essay sets (sourced from datasets like ASAP (Kaggle competition) or IELTS corpora) in systematically controlled ordering permutations: * High-to-Low Permutations: Sequencing exceptional essays followed by average and poor submissions. * Low-to-High Permutations: Sequencing weak essays followed by average and high-quality submissions. * Randomised Control Sequences: Evaluating essays in shuffled sequences across multiple independent runs. Quantitative analyses will track changes in Quadratic Weighted Kappa (QWK), or other measures across these permutations to determine the magnitude of contrast and anchoring effects.
Resources
Datasets: Publicly benchmarked essay corpora with validated human ground-truth scores (e.g., Automated Student Assessment Prize / ASAP dataset, and/or IELTS writing corpora). LLM Access: API access to proprietary models (OpenAI, Anthropic, Google) alongside local GPU instances (e.g., NVIDIA RTX/A100) for running the models. (this is partially sponsored)
Background
Url
Difficulty Level
Moderate
Ethical Approval
None
Number Of Students
1
Supervisor
Ian Tan
Keywords
llm, automated essay grading, sequential priming
Degrees
Bachelor of Science in Computing Science