• Home  
  • LLM proposal rankings matched human reviewers while costing hundreds of times less
- Technology

LLM proposal rankings matched human reviewers while costing hundreds of times less

A Scientific Reports study found that LLM-based pairwise comparisons produced proposal rankings broadly aligned with human reviewers at a major US neutron facility, while estimated review costs were 346 to 823 times lower for typical proposal pools.

Clean editorial photograph-style illustration for a research news website. One clear main subject: a single printed scientific research proposal resting on a desk beside a softly glowing laptop screen suggesting artificial intelligence analysis. Minimal modern laboratory-office background, shallow depth of field, restrained professional lighting, uncluttered composition, realistic paper texture. No people, no readable text, no letters, no numbers, no labels, no logos, no watermarks, no interface elements.

Large language models may be able to support scientific proposal selection at major research facilities by comparing submissions directly rather than assigning each proposal an isolated score. A new peer-reviewed study found that an LLM-based ranking system broadly tracked human reviewer rankings and did not perform significantly worse at identifying proposals associated with stronger publication output.

The study, published in Scientific Reports on 25 September 2026, analysed historical proposals from three beamlines at the Spallation Neutron Source at Oak Ridge National Laboratory in the United States. Instead of asking an AI system to score each proposal independently, the researchers made the model compare proposals in pairs and select which submission showed greater scientific merit.

Why proposal ranking is difficult

Large research facilities must allocate limited experimental time among many competing proposals. The Spallation Neutron Source receives more than 500 to 600 proposals in a typical run cycle. Each proposal is generally reviewed by no more than four experts, while individual reviewers can be assigned as many as 12 proposals.

This structure creates an important measurement problem. Reviewers do not necessarily evaluate the same proposals, and differences in scoring standards, expertise, fatigue and judgement can make scores difficult to compare across the entire pool. The researchers therefore tested whether pairwise comparisons could provide a more internally consistent basis for ranking submissions.

How the researchers tested the LLM approach

The researchers used historical general-user proposals from three representative neutron-scattering instruments: EQ-SANS, CNCS and POWGEN. The analysis covered 17 to 20 run cycles, depending on the beamline. Earlier records were excluded where proposal ratings were incomplete.

Proposal PDFs were converted to text before the model compared every possible pair of proposals within a run cycle. Gemini 2.5 Flash was instructed to assess scientific merit, summarise the two proposals, compare them, explain its reasoning and select a winner or tie. The resulting win-loss data were converted into rankings using a Bradley-Terry model, a statistical framework designed to estimate relative strength from head-to-head outcomes.

The researchers then compared the LLM rankings with historical human rankings using Spearman correlations. They also linked accepted proposals to subsequent publication records to test whether either ranking system was better at identifying proposals that later generated more publications.

LLM rankings broadly aligned with human rankings

Across the historical cycles, correlations between LLM and human rankings varied substantially. Spearman correlations ranged from approximately 0.2 to 0.8. After removing 10% of outliers, correlations were at least 0.5, according to the study.

The publication analysis provided an additional test because there is no definitive ground truth for the correct ranking of scientific proposals. Among accepted proposals, the researchers found no statistically significant difference between the LLM and human rankings in their ability to identify proposals associated with high publication potential.

This does not establish that the model identified the scientifically best proposals. Publication output is an imperfect proxy for scientific value, and rejected proposals cannot generate publications from experiments that were never performed at the facility. The result instead suggests that, under the study’s retrospective evaluation, the automated ranking was not detectably worse than the historical human ranking on this particular outcome.

Estimated review costs were hundreds of times lower

The researchers also compared estimated human labour costs with model inference costs. They estimated a human review at $54.90 per proposal and an LLM pairwise comparison at $0.0046.

For proposal pools in the typical range of 30 to 70 submissions, the estimated human cost was between 346 and 823 times the LLM cost. Expressed another way, the LLM pairwise approach was estimated to cost approximately 0.12% to 0.29% as much as the conventional human individual-scoring approach.

The difference is notable because pairwise comparison requires many more evaluations as the number of proposals grows. For N proposals, a complete comparison requires N(N-1)/2 pairs. Despite this quadratic growth, the low model cost kept the approach substantially cheaper within realistic proposal-pool sizes examined by the researchers.

AI could also help identify unusually similar proposals

The study also tested proposal embeddings, which represent documents as numerical vectors. Similarity between these vectors could be used as a low-cost screening tool to surface pairs of proposals that appear unusually similar.

The authors stress that similarity is not itself evidence of duplication or misconduct. Highly similar submissions can represent legitimate resubmissions, different researchers studying the same topic, related methods applied to different systems, or proposals that simply share technical vocabulary. Expert review remains necessary to interpret any flagged pair.

Important limitations

The findings come from a retrospective analysis at one major US research facility and only three of its beamlines. Results therefore cannot automatically be generalised to grant agencies, universities or other scientific facilities with different disciplines, proposal formats and review criteria.

The proposal texts themselves could not be made publicly available, which limits independent replication using the identical dataset. The study also relied on one principal LLM for pairwise judgement, and future models or prompting strategies may behave differently.

Prospective use introduces additional risks that were less important in this historical dataset. Applicants who know an LLM is involved could optimise wording for the model, and automated systems can show positional, length or authority biases. The authors propose safeguards including swapping proposal order, comparing multiple models, removing author and institutional cues, and retaining human committees as the final decision-makers.

What the findings could mean for research governance

The results support a role for LLMs as decision-support tools rather than autonomous funding or access authorities. Pairwise AI rankings could give review committees an additional consistent comparison across a large proposal pool, identify cases where automated and human assessments diverge, and reduce the manual burden of screening for overlapping submissions.

The economic findings are also potentially important for facilities that process hundreds of applications but operate with limited administrative capacity. However, lower computational cost does not resolve questions about accountability, confidentiality, bias or scientific judgement. Any practical deployment would need governance rules that define how model recommendations are audited and how much influence they receive in final allocation decisions.

Source Information

Study: LLMs can assist with proposal selection at large user facilities

Authors: Lijie Ding, Janell Thomson, Jon Taylor, Changwoo Do and colleagues

Journal: Scientific Reports

Published: 25 September 2026

DOI: 10.1038/s41598-026-60226-1

Study type: Retrospective comparison of LLM and human proposal rankings using historical research-facility proposals and publication records

Research Today is a South African digital publication that makes credible research easier to understand.

 

ResearchToday.bus@gmail.com

TERMS OF USE & PRIVACY POLICY

follow us