Medical artificial intelligence is often developed under an assumption that more sophisticated architecture will produce better diagnostic performance. A newly published benchmark study suggests that assumption deserves more scrutiny.
Researchers tested a dual-stream system called MedFuse across 12 standardised medical image classification tasks. Instead of building an elaborate fusion mechanism, the system simply combined local image features from a trainable convolutional neural network with global representations from a frozen vision foundation model.
The result was not a universal win. Performance varied by task, but simple feature concatenation was competitive across many datasets and was often more stable than adaptive weighting or attention-based alternatives.
Local detail and global context were deliberately separated
Published in Frontiers in Medicine on 29 September 2026, the study evaluated all 12 two-dimensional subsets of MedMNISTV2. The benchmark spans colon pathology, chest X-rays, pneumonia X-rays, retinal OCT, breast ultrasound, blood-cell microscopy, tissue microscopy, dermoscopy, retinal fundus images and CT-derived organ classification.
The architecture used two parallel branches. A CNN was trained on each target dataset to learn fine local texture and morphology, while a pretrained vision foundation model remained frozen and supplied a global semantic representation.
For the default configuration, ResNet-18 produced a 512-dimensional local feature vector and DINOv2-ViT-S/14 produced a 384-dimensional global vector. MedFuse concatenated them into an 896-dimensional representation before classification, without learned weighting, gating or cross-attention.
The researchers compared accuracy and area under the receiver operating characteristic curve, or AUC, and independently repeated experiments three times. They also tested larger DINOv2 models, different CNN and transformer backbones, and more elaborate fusion strategies.
Strong results depended on the medical task
On PathMNIST, the default MedFuse configuration achieved an AUC of 0.989 and accuracy of 0.901. The ResNet-18 baseline reached an AUC of 0.950 and accuracy of 0.815, illustrating a sizeable improvement on that histopathology task.
BloodMNIST also produced strong results. MedFuse reached an AUC of 0.997 and accuracy of 0.951, compared with 0.996 and 0.932 respectively for ResNet-18.
The pattern was not consistent everywhere. On BreastMNIST, the default fusion model recorded an AUC of 0.854 and accuracy of 0.803, below ResNet-18’s 0.878 AUC and 0.827 accuracy. ChestMNIST likewise showed that combining representations did not guarantee improvement.
This variation is central to the paper’s message. The value of combining local and global representations appears to depend on imaging modality, dataset size, class distribution and the visual scale of the features needed to distinguish categories.
Bigger foundation models were not reliably better
The researchers then changed the scale of the foundation model and CNN. The largest pairing, DINOv2-ViT-L/14 with ResNet-50, did not produce the best average performance.
On OrganCMNIST, for example, the smaller DINOv2-ViT-S/14 plus ResNet-18 combination achieved an AUC of 0.992 and accuracy of 0.898. The larger DINOv2-ViT-L/14 plus ResNet-50 pairing recorded an AUC of 0.912 and accuracy of 0.889.
Across the scale comparison, the authors found no clear positive relationship between parameter count and average AUC or accuracy. Compatibility between the local and global representations appeared more important than simply adding model capacity.
That finding has practical relevance because medical AI development frequently operates under constraints that differ from consumer-scale computer vision. High-quality labelled clinical images can be expensive to obtain, and larger models can increase training, memory and deployment costs without guaranteeing a proportional gain.
More sophisticated fusion also came with trade-offs
MedFuse’s simplest design merely concatenated the two feature vectors. The team compared this with adaptive weighting, which learns how much emphasis to give each stream, and attention-based fusion, which attempts to model relationships between the representations.
The more complex methods produced good results on selected tasks but were less consistent. Attention fusion achieved an AUC of 0.705 on ChestMNIST and 0.992 on OrganCMNIST, yet accuracy declined sharply on several other datasets.
On PathMNIST, attention fusion reached 0.748 accuracy, while the default concatenation configuration reached 0.901. On OrganSMNIST, attention fusion recorded 0.648 accuracy compared with 0.791 for default MedFuse.
Simple concatenation has costs of its own. It increases the dimensionality of the classifier input and does not explicitly model interactions between the CNN and foundation-model representations. Those relationships must instead be learned downstream.
Benchmark performance is not clinical validation
The results should not be read as evidence that MedFuse is ready for diagnostic use. All experiments relied on MedMNISTV2’s predefined internal training, validation and test splits.
The researchers did not test whether performance transfers to independent hospitals, patient populations, scanners or acquisition protocols. They also did not perform formal significance testing, stratified k-fold cross-validation or report 95% confidence intervals.
Those limitations matter because some differences between models were modest. Repeating each experiment three times provides an estimate of variability, but a higher mean score cannot automatically be treated as a statistically reliable improvement.
The study nevertheless offers a useful engineering lesson. When labelled medical data are limited, architectural complexity is not automatically synonymous with better performance. A relatively simple way of preserving complementary local and global features can be a strong benchmark against which more complicated systems should have to prove their value.
For healthcare systems considering future AI tools, that distinction is important. The clinically valuable model will not necessarily be the largest or most intricate one, but the one that remains accurate, reproducible, efficient and robust when it leaves a benchmark dataset and encounters real patients.
Source Information
Study Title: MedFuse: dual-stream fusion of convolutional and vision transformer-based features for enhanced medical image classification
Authors: Yajing Ren, Ling Hai, Zheng Gu and Wen Liu
Journal: Frontiers in Medicine
Year: 2026
Published: 29 September 2026
DOI: 10.3389/fmed.2026.1901033
Study design: Benchmark evaluation across all 12 two-dimensional MedMNISTV2 subsets, with three independent experimental runs per configuration








