Universities are increasingly placing artificial intelligence between students and course material, but choosing an educational chatbot is not as simple as asking which system gets the most questions right. A chatbot can answer correctly once and then abandon the correct answer when challenged. Another can initially fail but successfully correct itself. A third can appear accurate while producing unstable answers when the conversation changes.
A newly peer-reviewed study in Scientific Reports proposes a standardized way to capture those differences. Researchers at the University of Namur in Belgium developed an “AI Score” that combines initial accuracy, robustness, self-correction and unreliability into a single benchmark for educational chatbots. When they applied the method to six platforms, the resulting scores ranged from 60% to 100%.
The study, published on 28 September 2026, is particularly relevant because it focuses on chatbots supplied with actual course resources through retrieval-augmented generation, or RAG. In other words, the systems were not simply asked to answer from their general training. They were given material resembling the documents an institution might provide to a course-specific AI tutor.
Why a correct answer is only the beginning
The researchers argue that conventional accuracy measures can conceal behaviour that matters in teaching. A student does not necessarily ask one question, receive one answer and end the interaction. Students question explanations, express doubt and ask systems to reconsider. That makes consistency and the ability to recover from errors educationally important.
The proposed AI Score therefore combines four components. Initial performance measures whether the chatbot gives the correct answer at the first attempt. Robustness tests whether it maintains a correct answer when subsequently challenged. Self-correction ability captures whether an initially incorrect answer can be repaired through follow-up interaction. Lack of reliability acts as a penalty for unstable behaviour.
The weighting deliberately gives the greatest influence to getting the answer right immediately. Initial performance contributes 70% of the positive score, robustness 20% and self-correction 10%, while lack of reliability carries a 25% penalty. The researchers acknowledge that these weights are a design choice rather than a universal law. Changing them could change the ranking, which is one reason the paper treats the score as a reproducible framework rather than a permanent league table of AI products.
Six chatbots faced the same course-grounded test
The validation compared six platforms: ChatGPT, Copilot Studio, NotebookLM, Grok, Mistral and Amanote. The systems were evaluated under standardized conditions using retrieval-augmented generation and course-specific resources. The published study reports validation across three first-year courses covering optics, Law of Obligations, and Economic and Social History, giving the framework a test across scientific and humanities-oriented material rather than a single disciplinary setting.
The optics component provides a particularly clear illustration of why ordinary test accuracy was not enough. Twenty multiple-choice questions were drawn from a previous examination. The chatbots had access to the syllabus, annotated lecture slides and lecture audio transcripts. Students, by contrast, had completed the examination without the same open-book access, so the student results should not be interpreted as a fair head-to-head contest between humans and machines.
On the first attempt, NotebookLM answered all 20 questions correctly. ChatGPT, Grok and Amanote each scored 17.5 out of 20, Mistral scored 16.25, and Copilot Studio scored 13.75. The comparison group of 266 students averaged 8.17 out of 20, with a median of 8.25 and a standard deviation of 4.48.
Those raw results might make most of the chatbots look similarly strong. Five of the six were already at or above roughly 80% on the initial examination. The deeper benchmark, however, produced much wider separation once the systems had to demonstrate that their answers remained dependable through follow-up interaction.
The final scores spread from 60% to 100%
NotebookLM achieved an AI Score of 100%, receiving an A grade in the framework. Amanote followed at 98%, also graded A, while Grok scored 95% and received the same grade. ChatGPT scored 85% and Mistral 83%, both graded B. Copilot Studio scored 60%, which placed it in grade E under the study’s classification.
The component results help explain why the composite measure separates systems that initially looked relatively similar. NotebookLM recorded 50 for initial performance, 50 for robustness and 50 for self-correction, with no lack-of-reliability penalty. Amanote recorded 50, 45 and 50 on those three components, also with no reliability penalty. Grok recorded 48, 46 and 48, again without a reliability penalty.
ChatGPT recorded 43 for initial performance, 40 for robustness and 44 for self-correction, with no reliability penalty, producing its 85% composite score. Mistral recorded 42, 41 and 46, likewise without a reliability penalty, for 83%. Copilot Studio recorded 38 for initial performance, 29 for robustness and 39 for self-correction, but it also accumulated a lack-of-reliability value of 24. That penalty contributed to the much lower 60% final score.
This distinction is the study’s central contribution. A model that performs reasonably well on an initial multiple-choice test may still be a weaker educational tool if it can be persuaded away from correct answers or behaves inconsistently when a learner probes its reasoning. Conversely, the ability to recognize and repair an initial mistake can be valuable in a conversational learning environment.
The benchmark is not a ranking of all AI use
The numerical differences are striking, but they require careful interpretation. The study tested specific systems, configurations, course materials and prompts at a particular point in a fast-moving technology cycle. Model updates can change performance. A platform that performs well with one course’s retrieval system may behave differently with another subject, another document set or a different prompting architecture.
The test also focuses on technical response performance. It does not show that a chatbot with a higher AI Score causes students to learn more, improves examination results or produces better long-term understanding. Nor does it directly measure whether students find a system motivating, usable or pedagogically supportive. Those are separate questions requiring learner-centred and longitudinal research.
The student comparison deserves similar caution. The chatbots effectively completed an open-book assessment because they had access to the course materials, while the students did not. The finding that the systems outscored the student average therefore demonstrates the machines’ ability to retrieve and apply supplied information under the test conditions. It should not be read as evidence that the systems possess greater understanding than students.
A practical pre-deployment test for universities
For educators, the framework offers a practical shift in how AI tools can be evaluated. Institutions often face rapidly changing product claims and model versions. A standardized course-specific test allows a university or lecturer to assess the system using the material students will actually encounter rather than relying only on generic vendor benchmarks.
The approach also highlights why adversarial follow-up questions matter. In a classroom, a confident student may tell a chatbot that its correct answer is wrong. A reliable tutor should not simply agree. Testing robustness therefore approximates a real conversational risk that ordinary one-shot accuracy misses.
Similarly, self-correction is not merely a technical convenience. An educational system will inevitably make mistakes. What matters is whether it can respond productively when a learner asks it to reconsider. By treating recovery as a measurable component, the framework recognizes that educational reliability is dynamic rather than confined to the first response.
Important limitations remain
The authors explicitly note several limitations. The weighting of the four criteria is partly normative and could be adjusted for different goals. The benchmark relies on multiple-choice questions, which are useful for objective scoring but cannot capture the full quality, nuance or explanatory value of open-ended teaching responses.
The evaluation also cannot establish the educational impact of deploying these systems with students over time. A technically reliable chatbot might still encourage superficial learning, be poorly integrated into teaching, or fail to support particular groups of learners. Conversely, a system with a lower technical benchmark might still provide useful learning experiences in a carefully supervised context.
Future research could therefore combine the AI Score with qualitative assessment of response quality, repeated testing as underlying models change, evaluation of different metaprompts and RAG configurations, and longitudinal studies measuring actual student learning. Extending the approach beyond the three courses used in validation would also reveal how stable the rankings are across disciplines.
Reliability needs to be tested, not assumed
The study does not settle which chatbot universities should use. Instead, it offers a way to make that decision more empirical. Its most important message is that a high first-attempt accuracy rate can hide meaningful differences in what happens next.
As educational AI becomes more conversational, the quality of the second and third exchange may matter almost as much as the first. A useful educational assistant must not only know when it is right. It should resist being pushed away from a correct answer, recognize when it is wrong and behave consistently enough that students can understand the limits of the tool they are using.
Source Information
Study: Dhyne, M., Meurisse, J.-R., Dumortier, L. et al. “A standardized AI score to benchmark educational chatbots.”
Journal: Scientific Reports
Published: 28 September 2026
DOI: 10.1038/s41598-026-70214-0
Research institutions: University of Namur, Belgium
Funding: The work acknowledges support from the University of Namur through the PUNCh GenAI4Student project. Michaël Lobet is a Research Associate of the Fonds de la Recherche Scientifique – FNRS.









