TY - GEN
T1 - Examining Students’ Code Comprehension with LLMs in Block- and Text-Based Programming
AU - Zhang, Shan
AU - Earle-Randell, Toni V.
AU - Prasad, Priyadharshini Ganapathy
AU - Liu, Zifeng
AU - Shi, Yang
AU - Bhat, Suma
AU - Israel, Maya
AU - Botelho, Anthony F.
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2026/2/17
Y1 - 2026/2/17
N2 - Understanding how students reason about code is essential for providing tailored scaffolding in computer science (CS) education. Prior work has used think-aloud protocols with the Structure of the Observed Learning Outcomes (SOLO) taxonomy to examine students’ code comprehension and programming levels. However, analyzing such data is labor-intensive and requires expert judgment. Recent advances in large language models (LLMs) offer a promising avenue for scaling this analysis, though their reliability for fine-grained coding remains uncertain. To address this gap, our study investigates the extent to which GPT-5 and 4o can classify SOLO levels and identify code-comprehension strategies from think-aloud transcripts of 27 high-school students working on block-based and text-based tasks. Results show modest alignment with human ratings for SOLO, with one-shot prompting improving agreement over zero-shot, though distinctions between adjacent lower levels (e.g., Prestructural 1 vs. 2) remained difficult. Strategy detection demonstrated stronger performance, achieving accuracies of 75–77% (block) and 62–67% (text), particularly for surface-visible strategies such as ‘walkthroughs’, ‘control-structure identification’, and ‘pattern recognition’, but weaker for less frequent, abstract, meta-cognitive strategies such as ‘strategizing’ (planning an approach) or ‘thoroughness’ (systematically checking work). These findings highlight both the potential and the limitations of using GPT-5 and 4o to analyze think-aloud data. While this work represents an initial step, with plans to examine more models, our preliminary results indicate that a human-in-the-loop approach is essential to ensure reliability and interpretive depth. Future work will extend this evaluation to other LLMs to better understand their role in supporting instructional decision-making.
AB - Understanding how students reason about code is essential for providing tailored scaffolding in computer science (CS) education. Prior work has used think-aloud protocols with the Structure of the Observed Learning Outcomes (SOLO) taxonomy to examine students’ code comprehension and programming levels. However, analyzing such data is labor-intensive and requires expert judgment. Recent advances in large language models (LLMs) offer a promising avenue for scaling this analysis, though their reliability for fine-grained coding remains uncertain. To address this gap, our study investigates the extent to which GPT-5 and 4o can classify SOLO levels and identify code-comprehension strategies from think-aloud transcripts of 27 high-school students working on block-based and text-based tasks. Results show modest alignment with human ratings for SOLO, with one-shot prompting improving agreement over zero-shot, though distinctions between adjacent lower levels (e.g., Prestructural 1 vs. 2) remained difficult. Strategy detection demonstrated stronger performance, achieving accuracies of 75–77% (block) and 62–67% (text), particularly for surface-visible strategies such as ‘walkthroughs’, ‘control-structure identification’, and ‘pattern recognition’, but weaker for less frequent, abstract, meta-cognitive strategies such as ‘strategizing’ (planning an approach) or ‘thoroughness’ (systematically checking work). These findings highlight both the potential and the limitations of using GPT-5 and 4o to analyze think-aloud data. While this work represents an initial step, with plans to examine more models, our preliminary results indicate that a human-in-the-loop approach is essential to ensure reliability and interpretive depth. Future work will extend this evaluation to other LLMs to better understand their role in supporting instructional decision-making.
UR - https://www.scopus.com/pages/publications/105033962127
UR - https://www.scopus.com/pages/publications/105033962127#tab=citedBy
U2 - 10.1145/3770761.3777328
DO - 10.1145/3770761.3777328
M3 - Conference contribution
AN - SCOPUS:105033962127
T3 - SIGCSE TS 2026 - Proceedings of the 57th ACM Technical Symposium on Computer Science Education V.2
SP - 1601
EP - 1602
BT - SIGCSE TS 2026 - Proceedings of the 57th ACM Technical Symposium on Computer Science Education V.2
PB - Association for Computing Machinery, Inc
T2 - 57th SIGCSE Technical Symposium on Computer Science Education, SIGCSE TS 2026
Y2 - 18 February 2026 through 21 February 2026
ER -