Understanding Learner Confusion Across Languages with Explainable AI
2025 (English)Independent thesis Advanced level (degree of Master (One Year)), 10 credits / 15 HE credits
Student thesis
Abstract [en]
This thesis investigates automated confusion detection in student discussion posts, focusing on both English and German data. Building on the pretrained EduDistilBERT model, its performance is evaluated in an exploratory cross-lingual application: German posts are translated into English through a machine translation pipeline and then classified using the English-trained model. To address the opacity of predictions, LIME is employed to analyze feature attributions and gain insights into decision-making. Results show that the model achieves high recall in both languages, with systematic differences in precision. In English, confusion is frequently overpredicted, while in German, exploratory analysis suggests that polite or formal expressions are often misclassified as confusion, and subtle indicators are overlooked. LIME analyses reveal language-specific lexical cues, such as pronouns, task-related words, and politeness markers, as key drivers of predictions. The study highlights the feasibility of applying educational NLP models across languages via translation, while emphasizing that findings on German data remain exploratory due to the smaller dataset size.
Place, publisher, year, edition, pages
2025. , p. 30
Keywords [en]
Confusion Detection, Online Learning, Explainable AI (XAI), LIME, Educational NLP, Cross-lingual Application, Machine Translation
National Category
Natural Language Processing
Identifiers
URN: urn:nbn:se:his:diva-25890OAI: oai:DiVA.org:his-25890DiVA, id: diva2:2003166
Subject / course
Informationsteknologi
Educational program
Data Science - Master’s Programme
Supervisors
Examiners
2025-10-032025-10-032025-10-03Bibliographically approved