• Home  
  • Confidence-based multiple choice changed how students expressed uncertainty, not their overall scores
- Education

Confidence-based multiple choice changed how students expressed uncertainty, not their overall scores

A study of optometry students found confidence-based multiple choice exposed uncertainty and reflection without significantly changing overall test scores.

Multiple-choice answer sheet with small counters divided between answer options to represent confidence weighting.

Multiple-choice tests are built around certainty. A student is shown several options and must commit to one, even when their actual knowledge is closer to “I think it is A, but B is still plausible.” That forced choice makes assessment efficient, but it also hides something educators may care about: how certain a learner is about what they know.

A new peer-reviewed study in BMC Medical Education tested whether a different format could expose that uncertainty without substantially changing students’ grades. Thirty-four first-year optometry students completed the same 50 questions using both a conventional single best answer format and a confidence-based format that allowed them to spread points across several answers when unsure.

The central result was reassuring for educators considering such an approach. Median scores were very similar. Students scored 60.00% in the conventional format and 61.25% in the confidence-based format, a median difference of just 1.25 percentage points that was not statistically significant. Yet the alternative format changed what students revealed about uncertainty, risk and reflection while answering.

A test that lets uncertainty become visible

Conventional single best answer tests remain popular because they are practical. They can cover broad content efficiently, are straightforward to administer and mark, and can be designed to test anything from recall to complex reasoning. Their weakness is that the final answer does not show how the student arrived there.

With four options, a student who has no idea can still guess correctly. A student who has confidently eliminated two wrong options but cannot distinguish between the remaining two is treated exactly like someone who is completely certain if both ultimately select the same answer.

The confidence-based answer test used in this study tried to preserve the familiar multiple-choice structure while recording that hidden uncertainty. For every question, students had four points to distribute. A confident student could place all four on one answer. Someone torn between two options could divide the points 2:2, or use a 3:1 split when leaning more strongly toward one. They could also divide points 2:1:1 or, when completely uncertain, 1:1:1:1.

The idea is not simply to make tests easier. Spreading points reduces the reward available from being fully correct, but it also allows partial knowledge to receive partial credit. That makes the pattern of point allocation informative in its own right.

How the researchers compared the two systems

The researchers invited 81 students enrolled in a first-year optometry unit at Deakin University in Australia. Across the trimester, students completed five summative assessments containing 10 multiple-choice questions each, giving 50 questions in total. Thirty-four students attended every relevant test and consented to having all their responses included in the performance analysis.

Crucially, this was not a comparison between two different sets of questions. Each participant answered the same questions in the same order under both scoring paradigms. The confidence-based response was completed electronically while the conventional single best answer response was recorded on paper.

The researchers compared paired scores, used Bland-Altman analysis to examine agreement, and tested whether score differences varied between stronger and weaker students. They also examined whether students who spread their points more frequently gained or lost more under the confidence-based system.

A separate anonymous exit survey broadened the evidence beyond the 34 students with complete performance data. Sixty-six students completed four Likert-scale questions about the testing formats, while 21 supplied free-text comments. Those comments were analysed thematically by two researchers, with the resulting themes reviewed by the wider research team.

Students used the option to hedge, but not most of the time

The point-spreading feature was used on 30% of questions. On the other 70%, students put all four points on a single answer, effectively behaving as they would in a conventional test.

When students did express uncertainty, the simplest hedge was the most common. A 50:50 split between two answers occurred on 16.3% of questions, followed by a 75:25 split on 9.8%. Three-way 50:25:25 allocations accounted for 3.2%, while an even four-way split appeared on only 1.2%.

That pattern is useful because it shows that confidence-based testing did not turn every question into an exercise in strategic point allocation. Most answers remained decisive. The extra information appeared primarily on the subset of questions where students perceived meaningful uncertainty.

Overall grades barely moved

The conventional test produced a median score of 60.00% with an interquartile range of 11.75 percentage points. Under confidence-based scoring, the median was 61.25% with an interquartile range of 11.41. The median difference was 1.25 percentage points and was not statistically significant, p = 0.31.

Just as importantly, the alternative format did not appear to systematically favour high-performing or low-performing students. The Bland-Altman analysis found no relationship between the difference in scores and the students’ average performance across the two methods.

Nor did frequent point spreading create an obvious score advantage. The relationship between how often a student used confidence weighting and the difference between their two test scores was extremely weak, R² = 0.02, p = 0.46.

There was, however, a different relationship worth noticing. Students who spread points more often tended to have lower overall test scores under both systems. The reported R² values were 0.21 for conventional scoring and 0.26 for confidence-based scoring. That does not mean hedging caused poorer performance. A more plausible reading is that students with weaker command of the material encountered uncertainty more often and therefore used the feature more frequently.

Partial knowledge did not automatically produce extra marks

The question-level results show why confidence weighting did not simply inflate grades. When students used a 75:25 split, their average confidence-based return was 0.45 points per question on a one-point scale, compared with 0.48 under the conventional scoring of those same questions. For 50:50 splits, the corresponding figures were 0.35 and 0.39. For 50:25:25 splits, they were 0.31 and 0.32.

The one exception was complete uncertainty. On questions where students divided their confidence equally across all four options, the confidence-based system guaranteed 0.25 points. Their conventional single-answer choices on those same questions averaged 0.23.

This is a small but revealing distinction. Confidence-based scoring can recognise that a student genuinely does not know the answer without pretending that a lucky guess demonstrates full knowledge. At the same time, students who hedge between plausible answers sacrifice the possibility of receiving a full mark unless they commit most of their points to the correct option.

The biggest change was psychological, not numerical

The exit survey suggests that making uncertainty explicit altered the experience of taking the test more than it altered the final score. Of the 66 survey respondents, 48, or 72%, agreed or strongly agreed that assigning relative weights made them more doubtful of their answers.

That doubt was not uniformly viewed as harmful. Forty-one percent agreed that assigning weights was beneficial to their learning and knowledge, compared with 23% who disagreed. Yet only 32% thought the method helped them score higher, while 23% disagreed with that proposition.

Students also failed to converge on a preferred format. Thirty-three percent said they preferred the confidence-based model, while 38% preferred the conventional single-answer approach. The median response was neutral.

The researchers checked whether the four survey questions were behaving coherently. The overall Kaiser-Meyer-Olkin measure was 0.78 and Bartlett’s test of sphericity was significant at p < 0.01. A single factor with an eigenvalue of 2.64 emerged, and all four items had extraction coefficients above 0.3. In other words, the short survey appeared to capture a common underlying response to the assessment format rather than four unrelated reactions.

Uncertainty became part of the learning conversation

The free-text responses help explain the mixed survey results. Four broad themes emerged: confidence, risk-taking, motivation during assessment and the learning journey.

Some students found that being allowed to split points made them second-guess answers they might otherwise have selected quickly. Others used the feature as a deliberate representation of what they knew, especially when they could narrow a question to two plausible answers. The same mechanism could therefore feel either distracting or intellectually honest depending on the student and the question.

Risk preference also mattered. Some students preferred to commit to one answer because they valued the chance of a full mark. Others saw a split as a rational way to protect themselves when uncertainty was genuine. This means confidence-based assessment does not merely measure factual knowledge. It can also reveal how learners manage uncertainty when consequences are attached to their confidence.

For educators, that information could be valuable. Two students with the same conventional score may have very different knowledge profiles. One may answer confidently and consistently, while another reaches the same score through a mixture of partial knowledge and successful guesses. A confidence-based response pattern offers a way to distinguish those cases without replacing the multiple-choice format entirely.

Why this does not prove confidence-based testing is better

The study supports equivalence more strongly than superiority. It found no meaningful overall score difference and no obvious performance bias across ability levels, but it did not demonstrate that confidence-based testing produces better long-term learning, more accurate clinical reasoning or more reliable high-stakes decisions.

The performance sample was also small. Only 34 students had complete paired data across all five tests, all came from a single first-year optometry unit, and the work was conducted at one Australian university. The survey included more students, but the qualitative component contained free-text comments from only 21 respondents.

Students answered each question using both formats in parallel rather than being randomly assigned to one format or the other. That was useful for direct within-person comparison, but exposure to one response process could influence the other. The electronic confidence-based format and paper conventional format also differed in delivery mode, although students were already familiar with both.

There is another practical consideration. A testing method that makes 72% of surveyed students feel more doubtful could be pedagogically useful if that doubt triggers reflection, but it could also increase cognitive burden or test anxiety for some learners. The current study was not designed to determine when productive uncertainty becomes counterproductive.

Assessment can record more than right or wrong

The appeal of conventional multiple choice is that it compresses a complicated cognitive process into a clean outcome. That efficiency is also what information gets lost. A correct answer cannot tell an educator whether the learner knew it immediately, eliminated three alternatives through reasoning, or guessed successfully.

This study suggests that some of that missing information can be recovered without radically changing overall marks. Students used confidence weighting selectively, scores remained closely aligned with conventional testing, and the added choice surfaced uncertainty that would otherwise have remained invisible.

For formative assessment, that may be the more important finding. Confidence-based questions could help identify where a student is correct but fragile in their knowledge, or wrong but already able to eliminate implausible alternatives. Those are different learning problems even when a conventional answer sheet treats them identically.

The next step is larger and more diverse testing. Researchers would need to establish whether confidence patterns predict later retention, clinical decision-making or misconceptions, and whether the approach remains fair across disciplines, student groups and high-stakes settings. For now, the evidence indicates that allowing students to say how unsure they are can change the information an assessment captures without necessarily changing the grade it produces.

Source Information

Study: A mixed methods evaluation of the effect of confidence-based versus single best answer multiple-choice questions on student performance and the learning journey

Authors: Luke X. Chong, Nick Hockley, Ryan J. Wood-Bradley and James A. Armitage

Journal: BMC Medical Education

Published: 26 September 2026

DOI: 10.1186/s12909-026-10458-6

Research Today is a South African digital publication that makes credible research easier to understand.

 

ResearchToday.bus@gmail.com

TERMS OF USE & PRIVACY POLICY

follow us