Images Versus Videos in Contrast-Enhanced Ultrasound for Computer-Aided Diagnosis
At a glance
- الاستشهادات
- 0
- المراجع
- 89
- Comments
- 0
Abstract
The background of the article refers to the diagnosis of focal liver lesions (FLLs) through contrast-enhanced ultrasound (CEUS) based on the integration of spatial and temporal information. Traditional computer-aided diagnosis (CAD) systems predominantly rely on static images, which limits the characterization of lesion dynamics. This study aims to assess the effectiveness of Transformer-based architectures in enhancing CAD performance within the realm of liver pathology. The methodology involved a systematic comparison of deep learning models for the analysis of CEUS images and videos. For image-based classification, a Hybrid Transformer Neural Network (HTNN) was employed. It combines Vision Transformer (ViT) modules with lightweight convolutional features. For video-based tasks, we evaluated a custom spatio-temporal Convolutional Neural Network (CNN), a CNN with Long Short-Term Memory (LSTM), and a Video Vision Transformer (ViViT). The experimental results show that the HTNN achieved an outstanding accuracy of 97.77% in classifying various types of FLLs, although it required manual selection of the region of interest (ROI). The video-based models produced accuracies of 83%, 88%, and 88%, respectively, without the need for ROI selection. In conclusion, the findings indicate that Transformer-based models exhibit high accuracy in CEUS-based liver diagnosis. This study highlights the potential of attention mechanisms to identify subtle inter-class differences, thereby reducing the reliance on manual intervention.
Publication details
- DOI
- 10.3390/s25196247
- OpenAlex
- W4415008643
- Document type
- article
- Language
- EN
- Source
- Sensors
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.