A standing long jump looks simple: crouch, swing, take off and land. Yet the distance recorded at the end compresses a complex sequence of coordinated movements into a single number. New research suggests that an ordinary side-view video can recover much more of that technique, and can predict jump distance substantially better than body measurements alone.
Researchers in China developed a phase-aware computer-vision system called KSHPV to analyse standing long jumps from monocular video. In testing across 204 participants and 219 analysed trials, a random-forest model using the system’s movement indicators predicted jump distance with a mean absolute error of 17.25 ± 2.67 cm and an R² of 0.665 ± 0.068. An anthropometrics-only baseline produced a much larger error of 28.92 ± 2.43 cm and an R² of just 0.160 ± 0.091.
That difference matters because standing long jump tests are widely used in physical education and fitness assessment, but conventional testing usually tells a participant only how far they jumped. It provides little explanation of why one attempt was better than another or which part of the movement could be improved.
Turning one jump into a sequence of measurable phases
The study, published in Scientific Reports on 3 October 2026, was designed around a practical constraint. Instead of requiring a multi-camera motion-capture laboratory, force plates or wearable sensors, the researchers used monocular side-view video. Their goal was to extract interpretable movement features from footage that could plausibly be collected in routine teaching and fitness settings.
The KSHPV pipeline first detects the athlete and estimates a two-dimensional skeleton across the video frames. It then identifies three key events: take-off, the peak of flight and landing. These events allow different jumps to be aligned by movement phase rather than compared frame by frame at arbitrary points in time.
From the aligned trajectories, the system constructs indicators describing countermovement preparation, propulsion, coordination between the upper and lower limbs, flight timing and trunk control during landing. A confidence-aware quality-control step is used to reduce the influence of unreliable pose estimates, an important consideration when field video contains imperfect detections or missing joint positions.
The researchers evaluated the resulting features with participant-grouped cross-validation. Grouping by participant is important because it reduces the risk that nearly identical movement characteristics from the same person appear in both training and test data, which could make predictive performance look better than it really is.
Movement information added far more than body measurements
The main comparison was striking. The phase-aware random-forest model reached a mean absolute error of 17.25 cm, compared with 28.92 cm for the anthropometrics-only baseline. In proportional terms, that is about a 40% reduction in mean absolute prediction error. The model also explained far more of the variation in jump distance, with mean cross-validated R² rising from 0.160 to 0.665.
The researchers then tested which parts of the method were responsible for the improvement. Phase alignment increased explanatory power by ΔR² = 0.289. In other words, putting jumps onto a common take-off, flight and landing timeline contributed substantially to the model’s ability to account for performance differences.
Coordination features supplied a smaller but consistent additional improvement of ΔR² = 0.012. That result is useful because it separates two ideas that are sometimes blended together in movement-analysis systems. Simply extracting more joint positions is not necessarily enough. Organising those measurements around meaningful phases appears to be particularly important.
The authors also examined robustness across gender and body-mass-index strata and under simulated noise and missing data. These tests were intended to establish where monocular pose-based assessment remains useful and where degraded video or incomplete skeletal tracking begins to undermine measurement reliability.
Why phase alignment changes the analysis
Two people can complete the same jump in different amounts of time. Even the same person may reach peak flight or initiate landing preparation at slightly different moments across attempts. If raw video frames are compared directly, one athlete’s propulsion phase may be matched against another athlete’s early flight phase.
By locating key events first, KSHPV compares functionally similar portions of the movement. This gives the model a better chance of distinguishing a genuinely different technique from a simple timing shift. The large ΔR² attributed to phase alignment supports that design choice within this dataset.
The approach also preserves interpretability. A black-box model might predict a jump distance accurately without offering a useful explanation to a coach or student. Here, the inputs are organised around recognisable components of the jump, such as preparation, propulsion, coordination, flight and landing control. That creates a path from prediction toward technique-oriented feedback.
Potential value for physical education and field testing
The practical attraction of monocular video is accessibility. Schools, sports programmes and community fitness settings are far more likely to have access to a phone or ordinary camera than to laboratory motion-capture systems. A reliable video pipeline could therefore add movement-quality information to tests that currently record only distance.
For teaching, this could help shift feedback from broad instructions such as “jump harder” toward more specific observations about preparation, timing or coordination. For repeated assessment, phase-aware indicators could also show whether technique changes even when the final distance changes only slightly.
The results should not be interpreted as showing that a camera can replace expert coaching or laboratory biomechanics. The reported 17.25 cm mean absolute error is meaningful, and prediction of final distance is only one validation target. A system intended to prescribe technique changes would need strong evidence that its individual kinematic measurements remain accurate across different cameras, environments, populations and movement styles.
Important limitations
The study’s strongest results come from a specific dataset and recording arrangement. Two-dimensional monocular pose estimation cannot fully reconstruct three-dimensional movement, and side-view footage can lose information when joints overlap or motion occurs outside the image plane.
The analysed dataset contained 219 trials from 204 participants, so the results should be viewed as a promising methodological validation rather than proof of universal performance. Wider testing would be needed across ages, ability levels, body types, camera positions, lighting conditions and devices. Real-world school environments may introduce occlusion, inconsistent framing and background clutter beyond what is represented in a curated dataset.
The random-forest results are predictive rather than causal. A feature that helps predict a longer jump is not automatically a technique that will increase jump distance if deliberately changed. Biomechanical relationships can be interdependent, and an apparently favourable movement pattern may reflect underlying strength, experience or other characteristics rather than a directly trainable cause.
Finally, the system’s value depends on reliable pose estimation. The confidence-aware quality controls and simulated missingness analyses address this problem, but they do not eliminate it. Field deployment would require clear rules for rejecting footage when measurement confidence falls below an acceptable threshold.
A richer fitness test from ordinary video
The study demonstrates a broader shift in computer vision for human movement: from simply recognising an action toward measuring its internal structure. In the standing long jump, aligning video around take-off, peak flight and landing allowed interpretable movement indicators to explain substantially more performance variation than anthropometrics alone.
For now, the system is best understood as a research framework with potential for scalable field assessment. Its strongest contribution may be the evidence that timing-aware movement structure contains useful information that a final distance score misses. If future validation confirms the measurements across more diverse real-world settings, a simple side-view recording could become a practical bridge between mass fitness testing and more informative biomechanical feedback.
Source Information
Study: Liu, Y., Kuang, G., Li, S. et al. “Phase-aware kinematic measurement of standing long jump based on human posture vision.”
Journal: Scientific Reports
Published: 3 October 2026
DOI: 10.1038/s41598-026-69788-6








