DOCTORADO EN INGENIERÍA - ÉNFASIS EN CIENCIAS DE LA COMPUTACIÓN · 2025
IMProB-It : automatic feedback model for iterative programming tasks in introductory programming courses
The growing interest in integrating artificial intelligence (AI) in education has spurred the development of tools and platforms that support learning, especially in computer programming. While many of these tools are designed to provide summative assessments and general feedback, only a few are tailored to address programming tasks that involve iteration structures, and an even smaller number incorporate Gries’ theory into their design. As a result, few solutions effectively support students’ understanding of loops in introductory programming courses. Students often struggle with loop construction, finding it difficult to grasp the scope of a loop, which code segments will repeat, and how many times they will execute. Automated feedback for loop-based programming can play a key role in improving students’ comprehension of these concepts. However, factors such as the type of feedback provided and the intervention level are essential to its effectiveness. This thesis addresses these difficulties by proposing an automated feedback model to support students’ understanding of iterative programming. Grounded in program correctness theory and machine learning (ML), the model evaluates student code, identifies common errors, and delivers targeted feedback to explain mistakes made in loop-based tasks. This research contributes significantly to the field of automated feedback in programming. We developed a specialized dataset of programming tasks featuring "while" loops, annotated to capture typical student errors, including issues with loop initialization, termination, and state transformation. Using Gries’ loop programming theory, we built a detailed taxonomy to categorize and label these errors. Later, we trained ML models on this dataset to classify and predict errors in students’ "while" loop tasks. Then, we employed prompt engineering with OpenAI’s GPT-4 to generate automated feedback aligned with Gries’ theory, tailoring it to the errors detected by the ML classifier. Finally, we integrated these models into INGInious, a learning management system (LMS), through an API named IMProB-It, allowing students to receive specific feedback on programming tasks involving "while" loops. To evaluate IMProB-It, we conducted a quasi-experimental study with first-semester students in systems engineering and related fields. Divided into experimental and control groups, students in the experimental group received automated feedback on their solutions to "while" loop tasks. The results indicate that students who received specific feedback found it beneficial for understanding loop mechanics, as reflected in surveymeasured satisfaction levels. This research demonstrates the potential of ML-driven automated feedback to enhance the learning experience for novice programmers, addressing an essential need in computer science education.
Texto completo 261 páginas con texto de 268
Leer la tesis completa Ficha en el repositorio
Contenido
- ming task with a "while" loopp. 114
- Generating incorrect versionsp. 114
- Program specification definitionp. 106
- Generation of the Synthetic Datasetp. 115
- Final Dataset: Key Characteristics and Relevant Informationp. 115
- Chapter Summaryp. 23
- Defining the Scope of the Automated Feedback Modelp. 122
- Error classification model in programming tasks with iterationsp. 123
- Description of preprocessing techniques for datap. 123
- tasks with iterationsp. 124
- sification modelp. 126
- Performance evaluationp. 121
- Prompt engineering for automatic feedback generationp. 121
- feedback to providep. 136
- ationp. 95
- Chapter Summaryp. 23
- Definition of Resources and Tools for the APIp. 148
- Diagrams and Functionality of the Automated Feedback APIp. 150
- Construction of the Automated Feedback APIp. 152
- Development of API endpointsp. 152
- Test of API endpointsp. 154
- Integration of INGInious with Feedback API Endpointsp. 156
- Functional Testing of the Automated Feedback APIp. 157
- Description of Test Scenarios for the APIp. 158
- Design of Test Cases Covering All Identified Scenariosp. 147
- Creation of Test Cases in INGIniousp. 147
- INGInious Setupp. 160
- Execution of Tests and Analysis of Resultsp. 162
- Chapter Summaryp. 23
- Definition of Target Population and Study Variablesp. 178
- Work Plan for Evaluating the Automated Feedback Modelp. 177
- Test Administration and Target Participantsp. 182
- Schedule for Tests Applicationsp. 182
- Evaluation of the Automated Feedback Modelp. 183
- Data Storage Detailsp. 177
- Study Results Documentationp. 183
- Comparative Analysis of IMProB-It with Other Studies and Limitsp. 177
- 10.1.1 Addressed Challenges and Goalsp. 206
- CV Keras model architecturep. 130
- Comparison of Metrics and Loss by Modelp. 131
- Classification report FC keras modelp. 132
- Confusion matrix FC keras modelp. 133
- Error Detection in Unseen Datap. 135
- adequate feedbackp. 139
- tion and bad feedbackp. 145
- tion and bad feedbackp. 145
- tion and correct feedbackp. 145
- ImproB-It architecture diagramp. 150
- ImproB-It C4 diagramp. 153
- INGInious - Problem statement and specifications and requirementsp. 160
- INGInious - Solution with errors Cases Testp. 162
- INGInious - Solution with errorsp. 163
- INGInious Course Imperative programmingp. 164
- 8.12 Results model error classification and feedback - GPT4p. 168
- 8.13 Results model error classification and feedback - GPT4 for tasksp. 168
- lutionsp. 172
- Comparison of Responses (GE vs GC) - Type of feedbackp. 185
- INGInious - Solution with errors Cases Testp. 186
- Comparison of Responses (GE vs GC) - Additional informationp. 188
- Comparison of Responses (GE vs GC) - Specificityp. 189
- Comparison of Responses (GE vs GC) - Importance Feedbackp. 190
- Comparison of Responses (GE vs GC) - Commitment to learningp. 191
- GE and CG Averagep. 194
- Comparison of means between GE and CG with standard deviationp. 195
- Mann-Whitney U values per questionp. 198
- 9.10 Comparison GE and CG Criterionp. 200
- portp. 235
- student codep. 238
- Statement and Student Codep. 238
- D.1 Results model error classification and feedback - GPT3.5p. 250
- D.2 Results model error classification and feedback - GPT3.5p. 250
- D.3 Results model error classification and feedback - Geminip. 251
- D.4 Results model error classification and feedback - Geminip. 251
- D.5 Results model error classification and feedback - LLAmap. 252
- D.6 Results model error classification and feedback - LLAmap. 252
- Distribution of labels and their counts in the datasetp. 116
- Ranges of values tested for hyperparameters in different techniquesp. 127
- Hyperparametersp. 128
- (green) and lowest (orange) values for each metricp. 128
- Programming Exercises with while Loops (Model evaluation.)p. 134
- Test cases for scenario "Calculate the Factorial"p. 161
- Integrated rubric for evaluation of feedback and model classificationp. 163
- Exercises and Their Complexityp. 169
- Integrated rubric for evaluating feedback in the surveyp. 181
- Test cases for each test scenario (basic programming tasks with loop)p. 247
- ImproBIt evaluation survey resultsp. 268