• Home  
  • AI monitoring model detected online-learning drowsiness at 86.8% mAP and estimated screen distance within 2.12 cm
- Education

AI monitoring model detected online-learning drowsiness at 86.8% mAP and estimated screen distance within 2.12 cm

A lightweight AI system combined yawning detection and screen-distance estimation for real-time online learning monitoring, with important privacy and validation limits.

Student using a laptop in an online learning setting

Online learning has made it possible to teach at a scale and distance that would once have been difficult to imagine. It has also created a deceptively simple measurement problem: when a student is sitting in front of a screen, how much can an educational system actually infer about whether that student is alert, physically positioned well and ready to engage?

A new peer-reviewed study in Scientific Reports approaches that question from an engineering perspective. Rather than trying to interpret a student’s full facial expression, gaze pattern or emotional state, researchers developed a lightweight computer-vision framework focused on two narrower signals: yawning and eye-to-screen distance.

The resulting model, called YOLO26-P2-CBAM, achieved a mean average precision at an intersection-over-union threshold of 0.5, or mAP@0.5, of 86.8% on the researchers’ yawning dataset. Its screen-distance estimates had a mean absolute error of 2.12 cm, while average inference took about 12.5 milliseconds per frame. On the public SCB dataset, the model reached an mAP@0.5 of 63.2%, 4.2 percentage points above the baseline YOLO26 model.

Those figures make the system technically interesting, particularly for resource-constrained online-learning environments. They do not, however, mean that a webcam can reliably determine whether a student is learning. The study measures the performance of specific visual detection tasks. Engagement, comprehension and motivation remain much broader educational constructs.

A narrower way to monitor learning behaviour

Many attempts to monitor online learning rely on detailed facial-expression recognition, gaze estimation or combinations of multiple behavioural signals. These approaches can demand substantial computing power and can become difficult to deploy when many users must be processed simultaneously.

The researchers instead designed a framework around drowsiness-related behaviour and viewing distance. Yawning provides an observable signal that may accompany fatigue, while distance from the screen can indicate whether a learner has moved unusually close to or far from the device. Neither measure is a direct test of learning, but both can be detected visually without requiring a complex interpretation of a person’s emotional state.

The model starts from the YOLO26 object-detection architecture and introduces three main modifications. A high-resolution P2 detection head is intended to improve recognition of small visual targets. A Convolutional Block Attention Module, or CBAM, helps the network emphasize informative spatial and channel features. The researchers also incorporated Wise-IoU loss, which is designed to improve bounding-box regression.

The technical changes matter because yawning and eye-related features can occupy relatively small portions of a webcam frame. A model that is computationally light but loses those small features would have limited practical value.

How the system was tested

The evaluation combined a public dataset with two datasets constructed by the researchers. The public SCB dataset provided a broader benchmark for detection performance, while the self-constructed datasets were used for yawning detection and screen-distance estimation.

The researchers also built the model into a browser-server collaborative monitoring system. Webcam video could therefore be processed in a configuration closer to a deployable online-learning application rather than testing the neural network only as an isolated laboratory model.

Parallel processing and batch inference were introduced for multi-user scenarios. This is an important design choice because an educational monitoring system that performs well for one video stream may become impractical when dozens or hundreds of simultaneous learners are involved.

Performance was assessed with standard object-detection measures including precision, F1 score and mAP@0.5. For distance estimation, the researchers reported mean absolute error. Inference time provided a measure of whether the system could process frames quickly enough for real-time use.

The public benchmark showed a measurable gain

On the SCB dataset, YOLO26-P2-CBAM achieved an mAP@0.5 of 63.2%. That was 4.2 percentage points higher than the baseline YOLO26 architecture.

The model also recorded the highest precision and F1 score among the methods evaluated in the study. Precision reached 62.2%, while the F1 score was 60.9%. The F1 score is useful here because it balances precision with recall rather than rewarding a model simply for being conservative about which detections it accepts.

The result suggests that the architectural changes improved the balance between finding relevant targets and limiting incorrect detections. It is also a reminder that performance depends strongly on the dataset. A 63.2% mAP result on the public benchmark is meaningful improvement, but it is not near-perfect recognition.

Yawning was detected more accurately in the purpose-built dataset

Performance was substantially higher on the self-constructed yawning dataset. Here, the model achieved an mAP@0.5 of 86.8%.

The difference between this result and the public-dataset result is important. Models often perform differently when image composition, camera conditions, participant behaviour and annotation practices change. A strong score on a purpose-built dataset demonstrates that the method can work under the tested conditions, but it does not establish that the same accuracy will carry over to every laptop camera, lighting environment or student population.

Yawning itself also needs cautious interpretation. A yawn can accompany tiredness, but it is not proof that a student is disengaged or failing to learn. A monitoring system can detect a visible event more confidently than it can infer the psychological meaning of that event.

Screen distance was estimated within a few centimetres on average

The second part of the framework estimated the learner’s eye-to-screen distance. The reported mean absolute error was 2.12 cm.

Mean absolute error expresses the average magnitude of the model’s distance error without allowing positive and negative errors to cancel each other out. A value of 2.12 cm therefore indicates that, under the study’s test conditions, estimated viewing distance was typically only a few centimetres away from the reference measurement.

This could be useful for ergonomic prompts or for detecting large changes in a user’s position. Yet the study should not be read as establishing a medically optimal screen distance, nor does a distance estimate by itself say anything definitive about concentration.

Real-time performance was part of the design

The average inference time was approximately 12.5 milliseconds per frame. In practical terms, that is fast enough to support real-time processing in the tested implementation.

Speed matters because educational technology operates under constraints that benchmark accuracy alone can hide. A model may be highly accurate but unusable if it requires expensive hardware or takes too long to process each frame. The researchers explicitly targeted a lightweight architecture and supplemented it with parallel and batch-processing strategies for multiple users.

The study therefore contributes as much through system design as through raw detection scores. It asks whether a relatively focused set of visual tasks can be monitored without the computational burden of richer facial and behavioural analysis.

What the findings do and do not say about learning

The strongest interpretation is technical: the modified model improved detection performance over its baseline on the public benchmark, performed strongly on the study’s yawning dataset, estimated viewing distance with relatively small average error and processed frames quickly.

The educational interpretation is more limited. The research did not demonstrate that detecting yawns improves grades, comprehension, persistence or long-term learning. It also did not show that students who yawn are necessarily less engaged. A system can recognize behavioural proxies without validating the full chain from proxy to educational outcome.

That distinction is particularly important in automated education. Once a numerical indicator is available, institutions may be tempted to treat it as an objective measure of attention. But visible behaviour is context dependent. Students can look away while thinking, sit at different distances because of vision or accessibility needs, or yawn for reasons unrelated to learning.

Privacy and governance matter as much as accuracy

A webcam-based monitoring system also raises questions that model-performance metrics cannot resolve. Continuous video analysis can affect privacy, autonomy and the relationship between learners and educational institutions. Deployment would require clear rules about whether video is stored, where processing occurs, how long derived data are retained and who can access alerts or behavioural records.

False positives matter too. Even a technically capable system can create harm if an occasional yawn or unusual seating position is interpreted as misconduct, poor effort or lack of engagement. The safest use would be supportive rather than punitive, with automated signals treated as prompts for voluntary assistance rather than judgments about individual students.

Accessibility should also be considered before broad deployment. Camera angle, mobility differences, facial characteristics, assistive devices and environmental conditions could all influence visual detection. Performance should therefore be validated across diverse users and real educational settings rather than assumed from benchmark results.

Where the research needs to go next

The study establishes proof of technical feasibility, but field validation is the next critical step. Future work could test whether performance remains stable across different webcams, lighting conditions, network speeds, ages and physical learning environments.

Researchers could also examine whether alerts based on these signals produce any measurable educational benefit. A randomized evaluation, for example, could compare a supportive fatigue or ergonomics prompt with ordinary online learning and measure outcomes such as persistence, comfort, task performance and retention.

Equally important would be auditing performance across demographic and accessibility groups. Aggregate accuracy can conceal systematic errors concentrated among particular users.

A promising detector is not an automatic teacher

The attraction of the new framework is its restraint. Instead of claiming to read a student’s mind, it focuses on two visually measurable behaviours and attempts to detect them efficiently. The 86.8% mAP@0.5 on the yawning dataset, 2.12 cm mean absolute error for screen distance and roughly 12.5 ms inference time show that those tasks can be handled with useful speed and accuracy under the conditions tested.

Its larger educational value will depend on what happens after detection. If systems like this are used to offer optional breaks, ergonomic reminders or other low-stakes support, they may become useful components of online-learning environments. If behavioural proxies are instead treated as definitive measures of attention or effort, the technology could outrun the evidence supporting it.

The study is therefore best viewed as a step in lightweight educational computer vision, not as proof that automated monitoring can measure learning itself.

Source Information

Original study: Diao, L., Wu, W., Lang, H., Li, J. and Zhao, H. “Learning behavior monitoring through drowsiness and screen-distance recognition.” Scientific Reports (2026).

Published: 25 September 2026.

DOI: 10.1038/s41598-026-71506-1.

Study type: Computer-vision model development and experimental benchmark evaluation using the public SCB dataset and two self-constructed datasets for yawning detection and screen-distance estimation.

Journal: Scientific Reports, a peer-reviewed Nature Portfolio journal.

Research Today is a South African digital publication that makes credible research easier to understand.

 

ResearchToday.bus@gmail.com

TERMS OF USE & PRIVACY POLICY

follow us