Digital assessments promise to show teachers whether a pupil is improving, not simply how the pupil performed on one test. But a new systematic review suggests that the research supporting this promise is uneven. Across 106 studies of digital progress-monitoring tools used with pupils who have, or are at risk of, special educational needs, only 8% reported strong evidence of responsiveness: the ability to detect meaningful change over time.
The review, published on 9 October 2026 in Frontiers in Education, examined studies representing a reported cumulative total of 774,420 students. That large figure needs qualification: some studies reused the same participant samples, so it is not a count of 774,420 unique children. The researchers found much more evidence about whether assessments produce consistent scores or correspond with other measures than about whether their scores reliably track progress during instruction.
Why tracking change is different from measuring achievement
Schools use assessments for several different purposes. A screening test may identify children who need additional support. A diagnostic assessment may clarify particular difficulties. Progress monitoring asks a different question: after weeks of teaching or intervention, is this child actually learning at a faster rate?
Those functions demand different forms of evidence. A test can rank pupils consistently on a particular day yet be too noisy to identify a small improvement across several weeks. A tool may correlate well with a longer examination without producing sufficiently stable growth estimates to guide decisions about changing instruction. The review highlights the importance of not treating those measurement properties as interchangeable.
This matters in multi-tiered systems of support and response-to-intervention programmes, where teachers use repeated results to decide whether to continue, intensify or modify teaching. An attractive dashboard and automatically generated line graph cannot make an unreliable measure of change reliable.
How the review was conducted
Jannik Nitz, Estella Schubert, Silena Müller and Thomas Hennemann followed the PRISMA 2020 reporting framework and used a synthesis aligned with COSMIN, an established approach to evaluating measurement instruments. They searched seven databases and screened 4,319 records initially. After removing records published before 2000 and duplicates, 3,223 records underwent title and abstract screening. The team assessed 319 full texts, provisionally included 126, and removed another 20 during adjudication to reach 106 unique studies.
The researchers assessed four principal domains: reliability, or score consistency; validity, or whether scores measure what they are intended to measure; responsiveness, or sensitivity to learning change; and classification accuracy, or the ability to distinguish pupils above or below relevant decision thresholds.
They defined a strong numerical coefficient in advance as 0.80 or higher, although the meaning of such a threshold depends on the specific statistic and context. Their synthesis therefore provides a useful evidence map, not a single pooled estimate of the performance of every digital tool.
The evidence gap in numbers
Of the 106 studies, 91 (86%) reported some validity evidence, but only 33% reported strong validity coefficients. Reliability appeared in 71 studies (67%), with 55% of all studies reporting strong reliability coefficients.
Evidence directly related to decisions about progress was less common. 48 studies (45%) reported responsiveness, yet only 8% of the complete study set reported strong responsiveness coefficients. Classification accuracy was examined in 38 studies (36%), with 31% of the overall set reporting strong coefficients.
These percentages are proportions of studies reporting particular kinds of evidence. They do not mean that only 8% of digital assessment tools work, that 92% are inaccurate, or that only 8% of pupils improved. The concern is a shortage of sufficiently strong, reported evidence for the exact use case of monitoring change.
Too few measurements to estimate growth well
The researchers also looked at study design. 15 studies (14%) measured pupils only once, making it impossible to assess responsiveness. Another 16 studies (15%) used just two measurement occasions. Two points can show a difference, but they cannot establish how stable a learning trajectory is.
65 studies (61%) collected three or more measurements, the minimum needed in principle for slope-based analysis. Yet only 17 studies (16%) reached eight measurement occasions, a threshold often recommended for more stable estimates of growth in curriculum-based measurement. The median study used just three time points.
This helps explain the mismatch between the availability of one-time validation statistics and the shortage of dependable evidence about progress. Estimating change requires more than administering a digital test twice. Researchers need enough observations, appropriate time intervals, and analyses that separate learning from ordinary measurement noise.
Reading dominated the research
The review found a strong concentration in particular subjects. 68 studies (64%) focused on curriculum-based reading measurement, while 21 (20%) examined mathematics. Four addressed language progress monitoring and nine examined direct behaviour ratings or digital behavioural observation. The researchers found no eligible studies in their ecological momentary assessment and wearables category.
Reading had the most developed evidence base, but the coverage was not uniform even there. For mathematics, classification accuracy appeared in only four of 21 studies (19%). In digital behaviour measurement, classification accuracy was reported in two of nine studies (22%); responsiveness was also reported in two of nine, with no strong responsiveness coefficients identified in that group.
These gaps matter because a tool suitable for measuring reading fluency cannot automatically be assumed suitable for tracking mathematical learning or behavioural change. Different constructs require different measurement strategies, validation samples and decision thresholds.
A large literature, but concentrated geographically
The review included 75 studies from the United States (71%) and 24 from Germany (23%). The remaining seven came from Canada, the Netherlands, Finland, India, Jordan and New Zealand. That concentration limits how confidently results can be transferred to schools with different languages, curricula, resources and assessment traditions.
The publication types also varied. 67 studies (63%) were peer-reviewed journal articles, 21 (20%) were technical reports and 18 (17%) were dissertations or other academic documents. Including grey literature can reduce publication bias, but it also means that the 106 source studies should not all be described as individually peer-reviewed journal publications.
Additionally, 29 studies (27%) were flagged as borderline digital cases because administration remained partly paper-based while scoring or data integration was computerised. The authors tested whether such methodological choices affected the broad findings and reported that the overall pattern remained robust across sensitivity analyses.
What teachers and school leaders can do
The practical lesson is not to abandon digital progress monitoring. Well-designed tools can reduce scoring time, organise repeated observations and help teachers notice pupils who may need more support. But school leaders should ask vendors and assessment teams for evidence matched to the intended decision.
For example, if a product claims to identify whether a reading intervention is working after six weeks, the relevant evidence includes growth sensitivity, stability of slopes, appropriate testing intervals and accuracy around meaningful decision thresholds. A single high correlation with a standardised achievement test does not answer those questions.
Teachers should also interpret graphs alongside classroom work, pupil engagement, curriculum coverage and professional judgement. A flat line may reflect limited progress, but it could also reflect inconsistent test conditions, a poorly targeted measure or ordinary measurement variation. Likewise, a rising line should not be treated as proof that a particular intervention caused improvement without considering other explanations.
Limitations of the review
The study is a systematic review of published and accessible research, not a new classroom trial showing how a particular digital product performs. Its 0.80 threshold is a practical categorisation, not a universal pass-or-fail rule for every statistic. Study populations, instruments and definitions differed, so percentages describing the evidence base should not be read as estimates of average test accuracy.
The cumulative student count includes some overlapping samples. The literature search also had language and accessibility boundaries, potentially excluding relevant validation work from other national assessment systems. A lack of eligible evidence in one category does not prove that no functioning tool exists in practice. Finally, because some included instruments were partly paper-based, the findings cannot be attributed solely to digital technology.
What needs to change in research
The authors recommend making responsiveness a standard part of validation for any instrument marketed as a progress-monitoring tool. Future studies should collect enough repeated measurements to estimate growth reliably, report classification performance where schools use cut-offs to make decisions, and expand beyond the relatively mature reading literature.
Researchers also need to validate tools in more languages, countries and school systems, with clear disclosure when participant samples overlap across publications. The ultimate goal is not simply more data points on a dashboard. It is trustworthy information that helps teachers recognise when support is working and when a pupil needs something different.
Source Information
Original study: Psychometric properties of digital progress monitoring in special education: a COSMIN-guided systematic review.
Authors: Jannik Nitz, Estella Schubert, Silena Müller and Thomas Hennemann.
Journal: Frontiers in Education, volume 11, Digital Education.
Published: 9 October 2026.
Study type: Peer-reviewed systematic review using PRISMA 2020 and COSMIN-aligned synthesis.
Evidence base: 106 studies; reported cumulative sample of 774,420 students, with some overlapping samples.
DOI: 10.3389/feduc.2026.1936281.







