article Open access

Multi-scale patch embedding and distribution-aware transformer learning for explainable histopathology analysis

  • Results in Control and Optimization
  • Elsevier BV
Research footprint

At a glance

Citations
0
References
65
Comments
0
Paper overview

Öz

This study introduces Histo-AdaptiveViT, a model designed for patch-level multi-class histopathology classification, focusing on generating local evidence for pathology workflows rather than serving as a standalone diagnostic tool. It employs a three-stage CNN to create fine-, medium-, and coarse-scale feature maps, which are merged into a unified token sequence to capture various histopathological cues. The model uses adaptive modulation with distribution-aware techniques in Transformer blocks to handle class imbalance. It includes a token confidence gate to reduce background and artifact noise and employs FUSE tokens to gather subtype-specific evidence. Moreover, it utilizes a distribution-adaptive loss that prioritizes rare classes while maintaining performance on more frequent ones. Histo-AdaptiveViT demonstrates strong performance across four public cancer datasets, achieving Macro-F1 scores of 0.994 for colorectal, 0.983 for breast, 0.992 for lung, and 0.951 for oral cancer, along with high MCC and balanced accuracy, surpassing baseline CNN and ViT models. Analysis revealed better recall for minority and confusing subtypes. Statistical tests indicate that performance gains are most significant where baseline results are not fully optimized. At severe downsampling, Histo-AdaptiveViT improves the worst-case Macro-F1 from 0.764-0.919 to 0.789-0.943, while also achieving higher robustness scores. Sensitivity analysis indicates that the three-scale design is optimal for balancing performance and computational cost. While performance decreases slightly during cross-domain transfers, Histo-AdaptiveViT retains class structure and outperforms in tail classes. Efficiency profiling shows that our model is viable for practical use, featuring 20.5M parameters, 4.74 GFLOPs, and 1.17 ms GPU time, suggesting feasibility for batched patch-level WSI pipelines. However, deploying WSI end-to-end needs further validation in tissue detection, tiling, quality control, disk I/O, heatmap rendering, slide-level aggregation, and pathologist review.

Record transparency

Publication details

DOI
10.1016/j.rico.2026.100768
OpenAlex
W7165562525
Document type
article
Language
EN
Source
Results in Control and Optimization
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.