Generative AI has moved into university classrooms faster than researchers have been able to establish what it actually does to learning over time. A semester-long study of business informatics students now adds a useful counterweight to both the enthusiasm and the alarm: students learned substantially, but access to generative AI did not produce a statistically measurable advantage in knowledge gain over conventional instruction.
The result does not mean AI was irrelevant. Students using AI reported more reflective engagement with the technology, while forms of cognitive effort associated with building understanding were linked to stronger learning. The more important finding may therefore be that simply adding AI to a course is not the same thing as improving learning.
A semester rather than a single AI task
The study, published in the Journal of Computer Assisted Learning, followed 87 second-semester business informatics students aged 18 to 29 across a nine-week university theory phase. The students were already divided administratively into three parallel classes, allowing researchers to compare three different instructional conditions without rearranging the cohort for the experiment.
One class of 33 students used AI in a structured tutor-like format. A second class of 26 could use AI without that structured guidance. The 28 students in the control class completed the course without AI support. All three classes were taught by the same instructor with the same teaching materials, helping reduce differences unrelated to the intervention.
Data were collected at the beginning, middle and end of the course. Knowledge was tested with 14 multiple true-false items containing five decisions each, producing a maximum of 70 points before scores were converted to percentages. The researchers also tracked motivation, cognitive load, critical thinking and reflective use. Their analyses included repeated-measures ANOVA and ANCOVA, regression, moderation and an exploratory mediation model.
Learning improved, but the AI groups did not pull ahead
Among the 53 students with complete knowledge data at the relevant time points, scores improved significantly across the semester. The time effect was substantial, F(1,50) = 29.87, p < 0.001, with a partial eta-squared of 0.374.
Descriptively, the unguided AI group recorded the largest average knowledge gain at 16.57 percentage points. The tutor-AI group gained 11.33 points and the control group 9.05 points. Those figures might look like an AI advantage at first glance, but the statistical comparison did not support that conclusion. The interaction between time and instructional group was not significant, F(2,50) = 0.77, p = 0.469, and the between-condition effect was small, with Cohen’s f = 0.18.
That distinction matters. A larger numerical mean in a small sample does not automatically establish a reliable treatment effect. The variation within the groups was considerable, especially in the unguided AI and control conditions, and the study did not have enough evidence to conclude that the observed differences reflected a genuine AI-driven learning benefit rather than sampling variation.
AI also did not widen the achievement gap
The researchers also tested a concern sometimes described as a Matthew effect: whether students who begin with stronger academic backgrounds or more AI experience benefit disproportionately, causing existing performance gaps to widen.
Neither prior academic achievement nor previous AI experience significantly changed the rate of knowledge improvement. A Bayesian model produced BF01 = 8.70, which the authors interpreted as moderate to strong evidence in favour of no moderation effect. Students with previous AI experience did have higher overall knowledge scores during the semester, F(1,48) = 7.04, p = 0.011, but they did not improve faster.
That finding should still be treated cautiously because almost the entire sample already had AI experience. Only two participants did not, leaving too little variation for a strong test of how AI familiarity changes learning.
The effort students invested mattered more
The cognitive-load results provide a more nuanced picture of what was happening during the course. Germane cognitive load, the mental effort directed toward constructing and organising knowledge, increased significantly from the middle to the end of the semester, F(1,37) = 10.86, p = 0.002. This increase occurred across conditions rather than being unique to AI users.
More importantly, germane cognitive load at the end of the course was positively associated with knowledge gain, with a standardised coefficient of β = 0.51 and p = 0.001. In a model that also included extraneous cognitive load, the two variables explained 23% of the variation in knowledge gain.
Motivation, by contrast, remained broadly stable. There was no significant change from mid-semester to the end and no significant difference in the trajectory between the three instructional groups. Motivation at the middle measurement also failed to predict knowledge gain, explaining only about 1% of its variance in that analysis.
Reflection changed even when test scores did not
Critical thinking scores remained stable and did not differ significantly between the AI-supported and control conditions. Yet reflective use showed one of the clearest AI-related differences in the study.
Students in the AI-supported condition reported significantly higher reflective use than students in the control group, F(1,40) = 20.21, p < 0.001, with partial eta-squared of 0.336. On the five-point scale, reflective-use averages at the final measurement were 3.72 in the combined AI condition and 2.82 in the control condition.
This suggests that an educational technology can alter how students engage with their own learning without immediately producing a detectable advantage on a knowledge test. That is particularly relevant to universities deciding how to integrate generative AI. The useful question may be less about whether students are permitted to use AI and more about what intellectual work the course requires them to do while using it.
Why the result should not be read as a verdict on AI
The study has important limitations. The classes were pre-existing rather than randomly assigned, so unmeasured differences in peer dynamics, scheduling or classroom culture could have influenced the results. Attrition was also substantial and uneven. By the final measurement, missing data relative to the first wave reached 12.1% in the tutor-AI group, 25.9% in the control group and 38.1% in the unguided AI group.
The researchers used complete-case analyses, which reduced effective sample sizes and statistical power. The study was also confined to one cohort of business informatics students, and it was not preregistered. These constraints mean a failure to detect an AI advantage is not proof that no such advantage can exist in larger samples, other subjects or differently designed courses.
What the evidence does challenge is the assumption that access to a powerful tool automatically translates into better educational outcomes. Across this course, learning happened in every condition. AI changed some aspects of engagement, but the strongest signals around knowledge were tied to the cognitive work students performed rather than to the mere presence of the technology.
What this means for universities
For universities rapidly rewriting assessment policies and teaching practices around generative AI, the findings favour instructional design over technological novelty. A chatbot can provide explanations and feedback, but students still need tasks that require them to organise information, test understanding and reflect on the quality of what they receive.
The South African higher-education context makes that distinction especially important. Institutions operate with widely different resources, class sizes and levels of digital access. If the educational return comes primarily from thoughtful integration rather than simply providing AI access, then effective adoption may depend as much on lecturer support, assessment design and student AI literacy as on the sophistication of the model itself.
The study ultimately offers a more restrained way to frame the AI-in-education debate. Generative AI was neither a shortcut to dramatically better learning nor an obvious source of deteriorating performance in this cohort. Its educational value appeared conditional. The technology became one part of a broader learning process whose outcomes still depended on how students thought, reflected and engaged.
Source Information
Study Title: Generative AI and Learning Dynamics in Higher Education: A Longitudinal Empirical Study
Authors: Chrysanthi Melanou, Maik Beege and Martin Kimmig
Journal: Journal of Computer Assisted Learning
Year: 2026
DOI: 10.1002/jcal.70322







