REVIEW article

Front. Artif. Intell., 26 June 2026

Sec. Machine Learning and Artificial Intelligence

Volume 9 - 2026 | https://doi.org/10.3389/frai.2026.1840804

Deep learning for cardiovascular disease: a comprehensive review of detection and risk forecasting

  • School of Computer Science Engineering and Information Systems, Vellore Institute of Technology, Vellore, India

Abstract

The biggest health threat to the global population is cardiovascular disease (CVD), which afflicts almost one-third of the global population and causes considerable monetary and social losses. Risk analysis should be performed in a timely and appropriate manner to enhance clinical practice and preventive interventions. The emergence of advanced data modalities, such as wearable sensors, medical imaging, electronic health records (EHRs), and genomic platforms, has led to a paradigm shift in the holistic assessment of CVD risk through multimodal data integration. This systematic review is a methodological analysis of recent multimodal input deep-learning algorithms that enhance the early detection of CVD and risk-specific evaluation, following PRISMA 2020 guidelines across 69 studies published 20,122,025 based on 2,847 initial database records. We characterized the wide range of available data streams: longitudinal physiological measurements (ECG, HRV, and BP), echocardiogram data, cardiac MRI and CT, lab/demographic data, behavioral/environmental data, and genomic/proteomic data. Mid-level, early, late, and attention-based fusion methods are described in the context of deep neural networks, such as CNNs, RNNs, BiGRU with attention, and hybrid CNN-LSTM networks. Comparative studies showed dramatic improvements in predictive accuracy (often over 98%), strength to missing or noisy modalities, and access to real-time, individualized recommendations. The best-performing DEEP-CARDIO BiGRU-Attention model had 99.9 percent accuracy on Framingham and Statlog benchmarks. A systematic review of 28 studies by Grad-CAM and SHAP confirmed the dominance of each in imaging and structured-data tasks, respectively (Rahman et al., 2024). The federated explainable FL-LSTM model achieved 99% AUC across three ECG databanks with complete privacy protection. We end with a systematic reproducibility, federated learning, equitable AI, and regulatory translation roadmap.

1 Introduction

Cardiovascular diseases (CVDs) are a group of multiple disorders of the heart and blood vessels, such as coronary artery disease, stroke, heart failure, arrhythmias, and peripheral artery disease, which is the leading cause of death in the world, with an estimated 17.9 million deaths annually, approximating 32 percent of all deaths globally (Holt et al., 2023; Hathaway et al., 2021; Barbieri et al., 2020). In high-income countries, the burden of CVD has decreased with improved healthcare, whereas in low-to-middle income nations, there is a growing risk burden due to metabolic syndrome, urbanization, sedentary lifestyles, and pollution (Wang et al., 2021; Ordikhani et al., 2022) (Figure 1; Table 1).

Figure 1

Table 1

RegionCVD mortality (per 100,000)Most common subtypeNotable risk factor prevalence
North America180Ischaemic Heart DiseaseHypertension (35%), Obesity (41%)
Western Europe160Ischaemic Heart DiseaseHigh cholesterol (33%)
East Asia205StrokeSmoking (28%), Salt intake (43%)
South Asia260StrokeDiabetes (16%), Hypertension (41%)
Sub-Saharan Africa250Hypertensive Heart DiseasePhysical inactivity (42%), HIV comorbidity (12%)

Regional CVD mortality and key risk factors.

CVDs do not only have a human cost in terms of premature mortality—30 percent before the age of 70—but also have a significant economic cost in terms of direct medical expenses, lost productivity, and social dependency. The socio-demographic change, rising levels of obesity, sedentary lifestyles, poor diets, pollution, and tobacco use have further increased the number of at-risk populations in all age groups and on all continents, making CVD an urgent global public health concern (Bhagawati et al., 2023; Addissouky et al., 2024). The traditional risk models, namely the Framingham Risk Score and the ASCVD calculator, are based on the assumption of a linear relationship and lack temporal dynamics, limiting generalizability, especially in women, ethnic minorities, and individuals with multimorbidity (Yun et al., 2022; Wang et al., 2021; Ordikhani et al., 2022). Wearable sensors, EHR platforms, omics technologies, and IoT infrastructure are proliferating, revolutionizing CVD prediction by generating multimodal data streams that were previously unavailable to deep neural architectures (Sornalakshmi et al., 2024; Ahmad et al., 2021). The multimodal data landscape includes: (1) Clinical data—laboratory values, diagnoses, comorbidities, medications, and family history; (2) Physiological indicators—longitudinal ECG, HRV, blood pressure, and data from wearable monitors; (3) Medical imaging—echocardiograms, cardiac MRI, CT angiography, and vascular data from mammograms; (4) Genomics/omics—polygenic risk scores, transcriptomics, and methylation signatures; (5) Behavioral/environmental—lifestyle, activity, diet, and exposures sampled using smartphones and apps (Figure 2).

Figure 2

A case in point is systems like DEEP-CARDIO, which streamline the use of deep neural architectures for real-time ECG, blood pressure, pulse, and glucose, classifying conditions and providing interventions in real time. Hybrid neural systems combining CNN and LSTM have achieved predictive accuracies over 98%, compared to traditional systems. Decision trees and SVMs are best applied to structured tabular data, whereas ANNs are used to process mixed clinical data, CNNs to medical images, and RNNs (including LSTM and GRU) on the sequential physiological data (Gupta and Singh, 2023; Zhang et al., 2021; Krishnan et al., 2021; Reshan et al., 2023; Almulihi et al., 2022; Vincent et al., 2022; Yashudas et al., 2024).

1.1 Literature search strategy and PRISMA flow

This review is based on the PRISMA 2020 guidelines. Searches in PubMed, Scopus, IEEE Xplore, Web of Science, and Google Scholar (2015–2025) used Boolean combinations of: deep learning, CVD, multimodal data, ECG, medical imaging, risk prediction, fusion, EHR, wearable sensors, and explainability (Stahlschmidt et al., 2022; Gao et al., 2020). Inclusion criteria: (1) peer-reviewed articles with original DL/ML CVD models; (2) at least one of the following clinical/physiological/imaging/genomic modality; (3) quantitative measures of performance reported. Exclusion criteria: (1) No original model; (2) non-CVD outcomes; (3) lack of methodological detail; (4) non-English; (5) pre-2015 publications; (6) duplicate datasets without new methodology. The first search has yielded 2,847 records in the database and 218 additional sources. After eliminating 312 duplicates, 2,753 records were screened, with 2,341 being eliminated at the title/abstract stage. Out of 412 full-text articles evaluated, 343 were filtered out. Final synthesis: 69 studies, with 59% (n = 41) published 2024–2025 (Singh et al., 2024; Rahman et al., 2024) (Figure 3).

Figure 3

A contemporaneous systematic review of 65 ML studies (Banerjee et al., 2025) confirms that despite high benchmark accuracy, a significant translational gap persists—driven by dataset overreliance, lack of external validation, and black-box opacity—underscoring the necessity of the XAI and multimodal integration focus of the present review.

2 Multimodal data sources in cardiovascular risk stratification

The paradigm shift in predicting CVD is based on integrating five complementary data streams that encode different types of physiological information that cannot be captured adequately by any single data stream (Stahlschmidt et al., 2022; Gao et al., 2020; Zhang et al., 2023; Li et al., 2024). EHR data provide consistent, stable demographic bases; wearable signals capture real-time temporal changes; medical imaging provides spatial morphological details; genomics marks lifetime risk factors; and behavioral/environmental data capture lifetime and exposure determinants (Vincent et al., 2022; Yashudas et al., 2024; Holt et al., 2023). At the center of this paradigm is the compilation of very large, high-resolution databases such as UCI Heart, Framingham, MIMIC-III, PhysioNet MIT-BIH, and Cardiology DICOM, which capture various aspects of cardiovascular risk, including demographic and laboratory data, continuous physiological recordings, and advanced imaging (Figure 4; Tables 2, 3).

Figure 4

Table 2

DatasetModalitySize and class distributionPatient demographicsImaging/signal specsAnnotation standardMulti-center and harmonizationAccess
UCI Heart DiseaseClinical + Labs303 records; 54.4% disease positive, 45.6% negativeAge 29–77 yrs. (mean 54.4); 68% male, 32% female; mixed US population14 clinical/lab features; no imagingCardiologist-confirmed diagnosis labelsSingle-center (Cleveland Clinic); no harmonization appliedarchive.ics.uci.edu
Framingham Heart StudyEHR + Longitudinal4,238 records; ~15% CVD event positiveAge 32–70 yrs.; ~52% female; predominantly White US cohortLongitudinal clinical variable: SBP, DBP, cholesterol, glucose, BMI, smokingPhysician-adjudicated 10-year CVD event outcomesSingle community cohort (Framingham, MA); standardized exam protocols across waveskaggle.com/framingham
MIMIC-IIIEHR + Time-series60,000 + ICU admissions; multiclass (cardiac, respiratory, sepsis)Age 18–90 + yrs.; ~56% male; racially diverse US ICU populationTimestamped vitals (HR, BP, SpO2), labs, ICD-9 codes, waveforms at 125 HzICD-9 coded diagnoses; physician notesSingle institution (Beth Israel Deaconess, Boston); de-identified per HIPAA; no cross-site harmonizationphysionet.org
PhysioNet MIT-BIHECG48 half-hour recordings; 17 arrhythmia classesAge 23–89 yrs.; 22 male, 25 female; mixed ambulatory patients2-channel ambulatory ECG at 360 Hz; MLII + V5 leadsBeat-level annotation by 2 independent cardiologists; adjudicated consensus labelsSingle center; signal standardized to 360 Hz; baseline wander removedphysionet.org
Statlog HeartClinical variables270 records; 55.6% disease, 44.4% healthyAge and sex distribution similar to UCI; European cohort13 clinical variables (subset of UCI features)Binary classification labels; source cardiologist-verifiedMulti-source compilation (UCI, Hungarian, Swiss); no explicit harmonization protocol documentedUCI/Statlog
UK BiobankMRI + CT + Genomics + EHR30,000 + cardiac MRIs; 500,000 participant cohortAge 40–69 yrs. at recruitment; ~54% female; predominantly White BritishCardiac MRI: 1.5 T scanner; short-axis cine, long-axis views; standardized DICOM formatAutomated + expert-reviewed segmentation; phenotype definitions via ICD-10Multi-site UK (>20 assessment centers); centralized MRI acquisition protocol; batch correction appliedApplication required
EchoNet dynamicEchocardiography video10,030 apical 4-chamber echo videos; EF labelsStanford Hospital patients; diverse age/sex; no explicit race breakdown publishedGrayscale AVI video; 112 × 112 pixels; ~34 fps; variable clip lengthEF measured by 3 independent cardiologists; mean EF used as ground truthSingle center (Stanford); scanner-normalized preprocessing pipeline providedechonet.github.io
Cardiology DICOMImaging (ECHO, MRI)Variable; dataset-dependentNot standardized; varies by contributing institutionDICOM-format cardiac imaging across modalitiesVaries; no unified annotation standardMulti-source; no formal harmonization protocolkaggle.com
DatasetModalitySizeAccessReferencesKey Features
UCI Heart DiseaseClinical + Labs303 recordsarchive.ics.uci.eduZhang et al. (2021); Reshan et al. (2023); Almulihi et al. (2022)14 + features
Framingham StudyEHR + Demographic + Longitudinal4,238 recordswww.kaggle.com/framinghamHolt et al. (2023)Age, SBP, glucose, HDL, smoking
MIMIC-IIIEHR + Time-series Physiology60,000 + ICUphysionet.orgHathaway et al. (2021); Barbieri et al. (2020)ICD codes, labs, treatment
PhysioNet MIT-BIHECG (Wearable)48 recordingsphysionet.orgYashudas et al. (2024); Sumalatha et al. (2024)Arrhythmia annotations
Statlog HeartClinical variables270–303 recordsUCI/StatlogYashudas et al. (2024)Formatted for ML benchmarks
Cardiology DICOMImaging (ECHO, MRI)Varieskaggle.com/imaging-cardiacWong and Tse (2021)Cardiac imaging modalities
UK BiobankMRI + CT + Genomics + EHR30,000 + cardiac MRIsApplication requiredZhang et al. (2023)Multi-modal; largest open CVD dataset
EchoNet dynamicEchocardiography video10,000 + labelledFully openZhang et al. (2023)Gold standard for echo-based DL

Comprehensive cardiovascular dataset profiles—demographics, imaging specifications, annotation standards, and harmonization.

Table 3

DatasetModalitySize/samplesKey featuresAccessReferences
UCI Heart DiseaseDemographic/labs30314 + featuresUCI repositoryZhang et al. (2021); Reshan et al. (2023); Almulihi et al. (2022)
Framingham studyDem, labs, longitudinal4,238+Age, gender, SBP, glucose, HDL, smokingKaggleHolt et al. (2023)
MIMIC-IIIEHR, labs, codes60,000+Time-stamped EHR; ICD, labs, treatmentPhysioNetHathaway et al. (2021); Barbieri et al. (2020)
Statlog heartClinical variables270–303Similar to UCI, formatted for MLUCI/StatlogYashudas et al. (2024)

Structured cardiovascular data sources (with key features).

2.1 Clinical and structured data

A basic backbone of CVD risk modeling is provided by EHR and structured data (diagnoses, lab values, vital signs, medication use, procedural history, and family history) (Reshan et al., 2023; Bhagawati et al., 2023). The field is supported by reliable publicly available datasets such as UCI Heart Disease (303 records, 14 + features), Framingham Study (4,238 + records), and MIMIC-III (60,000 + ICU admissions) (Table 4).

Table 4

ParameterMeasurement/encodingMedical relevanceSource exampleReferences
AgeYears (integer)Baseline CVD risk; interacts with most risk factorsUCI Heart, FraminghamZhang et al. (2021); Holt et al. (2023)
Sex/genderMale/female (binary)Men with higher risk earlier; women with post-menopauseAll datasetsZhang et al. (2021); Holt et al. (2023); Wang et al. (2021)
BP (SBP/DBP)mmHg (continuous)Hypertension is a primary modifiable CVD risk factorUCI, FraminghamZhang et al. (2021); Holt et al. (2023)
Diabetes mellitusYes/NoMultiplies CVD risk 2-4xAll datasetsSelvarathi and Varadhaganapathy (2023); Wang et al. (2021)
Cholesterol (LDL, HDL, total)mg/dL or mmol/LKey predictor of atherosclerosis progressionUCI, FraminghamZhang et al. (2021); Holt et al. (2023)
family historyBinaryCaptures polygenic and monogenic inherited riskUCI Heart, FraminghamZhang et al. (2021); Holt et al. (2023)

Clinical parameters and medical relevance in CVD prediction.

2.2 Physiological signals and wearable data

Real-time temporal variations, which are not captured by clinical data alone, are provided by physiological signals, namely ECG (at-rest ambulatory), HRV, continuous blood pressure, glucose monitoring, and SpO2 (Vincent et al., 2022; Holt et al., 2023; Krittanawong et al., 2019; Haq et al., 2020). State-of-the-art neural networks, such as RNN, LSTM, GRU, BiLSTM, and BiGRU with attention, have been confirmed on large datasets and have demonstrated exceptional performance in detecting subtle arrhythmias and impending CVD events (Table 5).

Table 5

SignalKey featuresPhysiological significanceDataset(s)AdvantageLimitationReferences
ECGPQRST intervals, arrhythmia typeCardiac electrophysiology, arrhythmia riskMIT-BIH, DEEP-CARDIOHigh temporal resolutionNoise/electrode sensitivityYashudas et al. (2024); Sumalatha et al. (2024)
HRVRMSSD, LF/HF, SDNNAutonomic tone, stress markersMIMIC-III, PhysioNetNon-invasive autonomic markerExternal factors reduce specificityYashudas et al. (2024); Sumalatha et al. (2024)
BP (beat-to-beat)SBP/DBP, variability indexHypertension and vascular stiffnessDEEP-CARDIO, MIMIC-IIIKey CVD indicator via wearablesCalibration requiredSornalakshmi et al. (2024); Yashudas et al. (2024)
SpO2% saturation, variabilityOxygenation, sleep apnea riskMIMIC-III, DEEP-CARDIOSimple continuous monitoringLow specificity; altitude sensitiveSornalakshmi et al. (2024); Yashudas et al. (2024)
GlucoseLevel, postprandial excursionsDiabetes-CVD comorbidityIoT, DEEP-CARDIOStrong metabolic CVD signalRequires frequent monitoringSornalakshmi et al. (2024); Yashudas et al. (2024); Selvarathi and Varadhaganapathy (2023)

Physiological signals and deep learning approaches for CVD detection.

The BiGRU-Attention method of the DEEP-CARDIO system was a 99.9% accurate multiclass CVD detector that used simultaneous biosensor data, a significant improvement over classical machine learning or single sensor methods (Krittanawong et al., 2019). The ESMO-optimized EAWO-DNN, by incorporating ESMO optimization, increased network lifetime and diagnostic sensitivity by guaranteeing efficient power clustering in IoT systems (An et al., 2019) (Table 6).

Table 6

Imaging type/modelParameters/use caseMedical utility/strengthDataset/SourceContributionLimitationReferences
EchocardiogramChamber volumes, EF, wall motion, strainHeart failure, hypertrophy, structural analysisEchoNet DynamicFine-grained, interpretable risk markersInter-operator variabilityWehbe et al. (2023); Amal et al. (2024)
CCTA (CT Angiography)Coronary calcium (CAC), plaque, stenosisAtherosclerosis quantificationKaggle CCTA, PACSQuantifies atherosclerosis and valve diseaseHigh imaging costBarbieri et al. (2020); Cocianu et al. (2023)
Cardiac MRIFibrosis, volumes, perfusion, functional indicesMyocarditi, HF viability assessmentUK BiobankRadiomics extracts hidden predictive markersHigh storage requirementsSelvarathi and Varadhaganapathy (2023); Sekar et al. (2012)
Mammogram (BAC)Breast arterial calcification, calcium massPredicts coronary artery diseaseBAC datasets, Wang et al.AI BAC detection matches expert-level performanceLimited generalizabilitySelvarathi and Varadhaganapathy (2023); Yun et al. (2022)
CNN (2D/3D)Heart/artery segmentation, image classificationLearns spatial and anatomical structures efficientlyEchoNet, CMR datasetsDeep spatial learningRequires many labeled imagesSumalatha et al. (2024); Selvarathi and Varadhaganapathy (2023); Yun et al. (2022)
U-Net/DenseUNetPrecise tissue and lesion segmentationHigh-resolution ROI handlingEchoNet, CCTARobust segmentation in noisy imagingNeeds GPU resourcesZhang et al. (2023); Selvarathi and Varadhaganapathy (2023)
ResNet/inceptionPlaque and calcification detectionTransfer learning; excellent classificationKaggle, CCTADeeper networks generalize betterLarge model sizeZhang et al. (2023); Yun et al. (2022)
Autoencoder + CNNDimensionality reduction and BAC classificationLowers feature complexity; enhances performanceMammogram BAC datasetsImproves AUC with fewer featuresRisk of overfittingYun et al. (2022)
CNN (2D/3D)Efficient heart/artery segmentation and image classification of spatial and anatomical structuresLearns spatial and anatomical structures efficientlyEchoNet, CMR datasetsDeep spatial learningRequires many labeled imagesSumalatha et al. (2024); Selvarathi and Varadhaganapathy (2023); Yun et al. (2022); Wani et al. (2024)
Mammogram (BAC)Breast arterial calcification, calcium massPredicts coronary artery diseaseBAC datasets, Wang et al.AI BAC detection matches expert-level performanceLimited generalizabilitySelvarathi and Varadhaganapathy (2023); Yun et al. (2022); Alzahrani et al. (2025)

Overview of cardiovascular imaging techniques, AI models, and data resources.

2.3 Medical imaging and radiomics

The imaging-based cardiovascular biomarkers, echocardiography, cardiac magnetic resonance (CMR), coronary computed tomography angiography (CCTA), and mammograms of breast arterial calcification (BAC) allow the knowledge of the spatial, structural, and functional parameters of the cardiovascular system (Gupta and Singh, 2023; Holt et al., 2023; Zhang et al., 2023; Yun et al., 2022). These imaging modalities can give direct, quantifiable measurements of myocardial phenotypes, including chamber size, ejection fraction, vascular wall thickness, plaque burden, and calcification patterns.

Integration of genomics, proteomics, behavioral, and environmental data enables the identification of rare and complex causes, lifetime risk markers, inflammation markers, and social determinants of CVD (Krishnan et al., 2021; Wong and Tse, 2021; Yun et al., 2022). Major repositories such as the UK Biobank and dbGaP support this line of research.

2.4 Genomics, proteomics, and multi-omics

Integration of genomics, proteomics, behavioral, and environmental data enables the identification of rare and complex causes, lifetime risk markers, inflammatory markers, and social determinants of CVD (Krishnan et al., 2021; Wong and Tse, 2021; Yun et al., 2022). Major repositories such as the UK Biobank and dbGaP support this direction of research (Figure 5) (Tables 7, 8).

Figure 5

Table 7

TypeParameter/definitionPredictive roleKey dataset(s)
SNP/Genetic lociPolygenic/monogenic variant allelesStable, lifetime CVD riskUK Biobank, dbGaP (Hathaway et al., 2021)
DNA methylationCpG site status, methylation clocksAge-independent, epigenetic riskdbGaP, TCGA (Yun et al., 2022)
RNAseq/TranscriptomicsmRNA, lncRNA, miRNA, etc.Inflammatory at-risk phenotypeTCGA (Yun et al., 2022)
ProteomicsCardiac/vascular protein levelsBiomarker of ongoing diseaseFrontiers Cardiovascular Medicine (Mulani et al., 2025)

Multi-omics biomarkers in CVD risk assessment.

Table 8

Data typeTypical featuresExample datasetAccessReference
Demographics/EHRAge, sex, BP, cholesterol, comorbidities, ICD codesUCI Heart, Framingham, MIMIC-IIIPublic/KaggleZhang et al. (2021); Reshan et al. (2023); Almulihi et al. (2022)
Wearable signalsECG, HRV (RMSSD, LF/HF), BP, SpO2, glucoseMIT-BIH, PhysioNet, DEEP-CARDIOPhysioNetVincent et al. (2022); Yashudas et al. (2024); Holt et al. (2023)
Medical imagingECHO (EF, strain), cardiac MRI, CCTA (CAC, plaque)EchoNet Dynamic, UK BiobankOpen/applicationHathaway et al. (2021); Stahlschmidt et al. (2022)
Genomics/ProteomicsPolygenic risk scores, GWAS, SNPs, DNA methylationUK Biobank, dbGaP, TCGAApplicationKrishnan et al. (2021)
Behavioral/environmentalSteps, sleep, diet, stress, air pollutionApp-based monitoring platformsMobile appsYashudas et al. (2024); Holt et al. (2023)

Overview of major multimodal data types, features, and datasets.

The EAWO-DNN model shows that energy-sensitive architectures based on ESMO optimization can achieve accuracy beyond 98% and can ensure the life of an IoT network, making continuous real-time monitoring of CVDs clinically feasible (Sornalakshmi et al., 2024). The DEEP-CARDIO system builds upon this by combining streams of wearable biosensor data with EHR in a BiGRU-Attention pipeline and achieves 99.9% accuracy on proven benchmarks (Yashudas et al., 2024).

2.5 Preprocessing, augmentation, and multi-center data harmonization

Robust preprocessing and standardized augmentation are prerequisites for reproducible multimodal CVD modeling. Across the 69 reviewed studies, the following pipelines were consistently applied.

Structured/EHR Data: KNN imputation (k = 5, most common) or median imputation for skewed distributions was used to deal with missing values. Continuous variables (age, BP, cholesterol) were normalized to Z-scores, whereas binary/categorical variables were coded in a one-hot manner. Class imbalance is a widely used concept in datasets such as UCI Heart (54:46) and Framingham (average ratio of events 85:15), as well as in the training of neural networks. Class imbalance is widely observed in datasets such as UCI Heart (54:46) and Framingham (average ratio of events 85:15), as well as in the training of neural networks.

ECG/Physiological Signals: Raw ECG signals were bandpass filtered (0.540 Hz) to remove baseline wander and high-frequency noise, and R-peak detection was done using the Pan-Tompkins algorithm. The signals were grouped into fixed-length windows (usually 5–30 s) with half the window length. The augmentation strategies were: random amplitude scaling (±10%), Gaussian noise addition (SNR 2030 dB), time-warping, and lead dropout simulation to enhance model robustness to real-world signal degradation (Yashudas et al., 2024; Gao et al., 2020).

Medical Imaging: Images obtained by cardiac MRI and echocardiography were resized to a standardized size (224×224 or 112×112 pixels, depending on the architecture). Normalization of intensity was performed per scan, and min-max scaling was used. Standard augmentation was: random horizontal flipping, rotation (± 15 degrees), random cropping, brightness/contrast jitter, and elastic deformation of MRI volumes. In the case of CCTA, windowing (Hounsfield unit) was used [before CNN input (Zhang et al., 2023; Wehbe et al., 2023)] to window Hounsfield unit windowing (Hounsfield unit windowing).

Multi-Center Harmonization: Cross-institutional studies [MIMIC-III, UK Biobank, multi-site federated learning studies (Alasmari et al., 2025; Otoum et al., 2024)] harmonized through: (1) ComBat batch correction of imaging features, (2) feature-level z-score normalization per site, and then fused, (3) federated averaging (FedAvg) to prevent inter-institutional data leakage. Such studies without any protocols for explicit harmonization [Statlog, Cardiology DICOM]. This represents a limitation in reproducibility observed during our quality assessment.

Annotation Standards: The quality of annotations varied significantly across datasets. Multi-reader adjudication, reporting inter-rater reliability (Cohen, 0.80), was used on gold-standard datasets (MIT-BIH, EchoNet Dynamic). Administrative datasets (MIMIC-III) were based on ICD coding that is known to have miscoding rates of 5–15%. Reviewers discovered that 34 of 69 (49%) of those included studies did not report inter-annotator agreement measures, a crucial gap in clinical translation.

3 Deep learning fusion architectures for multimodal CVD risk prediction

It has been demonstrated that the introduction of integrated deep neural networks for multimodal CVD data has established a new gold standard in prediction analytics. New-generation models do not merely incorporate clinical and structured health records, but also include imaging, physiological sensor responses, and molecular profiles to develop a dynamic perspective of disease stratification that is beyond the reach of traditional single-modality models (Stahlschmidt et al., 2022; Gao et al., 2020) (Figure 6).

Figure 6

The underlying rationale is that modalities are statistically complementary: EHR data have stable baselines, wearable signals have temporal dynamics, imaging data have spatial morphology, and genomics data have lifetime risk markers. CNNs are used to analyze images; RNNs (LSTM, GRU) are used to analyze sequential data, whereas decision trees and SVMs are used to analyze structured, tabular data (Figure 7; Table 9).

Figure 7

Table 9

Model nameModalitiesArchitectureTop accuracyReferences
DEEP-CARDIOWearables + EHRBiGRU + Attention99.9%Yashudas et al. (2024)
HDNN hybridEHR + SignalCNN-LSTM98.86%Reshan et al. (2023)
EAWO-DNNIoT-Edge DataOptimized DNN>98%Sornalakshmi et al. (2024)
CNN + BiLSTMEcho + HRV2D CNN, BiLSTM97.5%Krishnan et al. (2021); Almulihi et al. (2022)
Transformer-basedMulti-sourceTransformer96–98%Browne et al. (2024)
Deep heartWearable + demographicLSTM-CNN87.6%Vincent et al. (2022)
Physics-guidedIoT + ClinicalPhysics-DL>95%Zhang et al. (2023)
Fusion-XAI cardiacEHR + ECG + clinicalMultimodal fusion + XAI97.8%Banerjee et al. (2025)

Recent deep learning models for CVD prediction.

A deep learning fusion architecture involves working at multiple integration levels. In the first type of fusion (called early fusion), raw feature vectors of all modalities are combined at the input layer; in the second type of fusion (called mid-level fusion), a modality-specific feature-vector is processed independently by separate subnetworks, and then combined with them at a hidden layer; in the third type of fusion (called late fusion), independent model outputs are combined through ensembling (Figure 8; Table 10).

Figure 8

Table 10

Fusion levelDescriptionExample architectureReference
EarlyRaw feature concatenation at the input layerCNN-LSTM with lab + ECGSornalakshmi et al. (2024); Zhang et al. (2023); An et al. (2019)
Mid-levelSubnetwork per modality, merge at the intermediate layerCNN + BiGRU with imaging + wearablesVincent et al. (2022); Krittanawong et al. (2019)
LateIndependent model outputs fused via ensemble methodsEnsemble of several DNNsZhang et al. (2023); An et al. (2019)
AttentionDynamic weighting based on model-learned importanceBiGRU with attention, DeepRiskVincent et al. (2022); Yashudas et al. (2024); An et al. (2019); Fathima and Fasla (2024)

Levels of multimodal data fusion in contemporary deep learning models.

3.1 Comprehensive architecture comparison

Table 11 lists the state-of-the-art models that have been produced by these fusion methods, including the DEEP-CARDIO model with BiGRU-Attention fusion, hybrid CNN-LSTM models, many of which commonly achieve prediction accuracies of 98–99.9%—a noticeable improvement over prior statistical methods.

Table 11

Model/architectureModalitiesSubnetwork typesDatasetsPerformanceBest forLimitationReferences
CNN + GRU hybridECG (1D) + EHR1D CNN, GRUIoT, DEEP-CARDIO97–98%Moderate heterogeneityModerate complexityKrishnan et al. (2021); Yashudas et al. (2024)
CNN + BiLSTMEcho + HRV2D CNN, BiLSTMEchoNet, Custom97.5%Image + signal fusionFeature alignment neededKrishnan et al. (2021); Almulihi et al. (2022)
Dual-branch ResNetClinical + CMRResNet, DNNMIMIC, UK Biobank96.2%High-res imaging + metadataHigh model complexityWehbe et al. (2023); Zhu et al. (2025)
EAWO-DNNSensor + EHRDNN + ESMOIoT, Custom>98%, AUC 0.99Low-power edge deploymentRequires embedded optimizationSornalakshmi et al. (2024)
HDNN hybridEHR + ECGCNN-LSTMCleveland, Multi-center98.86%Clinical + signal fusionHigher annotation needsReshan et al. (2023)
BiGRU-Attention (DEEP-CARDIO)Wearables + EHRBiGRU + AttentionFramingham, Statlog99.9%Complex multimodal integrationComputational costYashudas et al. (2024)
Synergized fusion-XAIEHR + ECG + ClinicalMultimodal Fusion NetworkCardiac Benchmark97.8%Fusion accuracy with clinical explainabilityRequires modality alignment[New]

Comparative analysis of multimodal deep learning architectures.

3.2 Attention mechanisms in CVD prediction

The BiGRU-Attention architecture performs better than hybrids based on CVD risk signals, since these are not stationary; a glucose spike may dominate prediction in one time window, and HRV anomalies in another (Yashudas et al., 2024; An et al., 2019). Attention mechanisms dynamically reweight modalities with respect to each patient state and cannot be done with a static fusion strategy (Figure 9; Table 12).

Figure 9

Table 12

Attention typeMechanismClinical benefitImplementationReference
Temporal attentionWeighs different time points in physiological signalsIdentifies critical cardiac eventsBiGRU-Attention modelsYashudas et al. (2024); An et al. (2019)
Feature attentionWeighs different clinical featuresHighlights the most predictive risk factorsSHAP-integrated attentionSaha et al. (2023); Dehghani et al. (2025)
Modality attentionWeighs different data modalitiesAdapts to available data sourcesCross-modal attention networksAn et al. (2019); Browne et al. (2024)
Spatial attentionWeighs different image regionsFocuses on pathological areas in imagingCNN with spatial attentionWehbe et al. (2023); Amal et al. (2024)

Attention mechanisms in CVD prediction.

3.3 Fusion strategy trade-off analysis

The strongest strategy in high-missingness settings is late fusion, since each unimodal model makes an independent prediction that can be combined even when some unimodal models receive no input. Attention fusion clearly increases interpretability through attribution of context-dri (Table 13).

Table 13

CriterionEarly fusionMid-level fusionLate fusionAttention fusion
Predictive accuracyModerateHighModerateVery high
Missing modality robustnessLowModerateHighHigh
Interpretability (XAI)LowModerateModerateHigh
Cross-modal feature learningHighHighLowHigh
Computational costLowMediumMediumHigh
Edge/IoT suitabilityHighModerateModerateLow

Fusion strategy trade-off matrix—operational strengths and clinical deployment suitability.

Complementary to these results, synergistic fusion modeling that integrates layers of XAI directly into multimodal pipelines spanning ECG, EHR, and clinical modalities has demonstrated that fusion accuracy and clinical interpretability need not be traded off, achieving about 97.8% accuracy and providing feature-level attribution across these modalities (Figure 10).

Figure 10

4 Validation, benchmarking, and clinical translation

Multimodal CVD AI models should be robustly validated with cross-validation using k-folds, bootstrapping, multi-center external holdouts, and calibration measures such as the Hosmer–Lemeshow test, Brier Score, and Harrell’s C-index. The excellence of multimodal architectures is exhibited in all performance aspects (Table 14).

Table 14

Model/approachAcc (%)Precision (%)Recall (%)F1 (%)AUCDataset(s)References
EAWO-DNN (IoT)98.998.898.798.80.99CloudSim, Real IoTSornalakshmi et al. (2024)
NSGA-II ensemble DL97.391.392.892.10.97UCI HeartGupta and Singh (2023)
Embedded FS + DNN98.697.899.398.30.983Kaggle, UCIZhang et al. (2021)
HDNN (CNN-LSTM)98.997.498.898.70.91Large Multi-centerReshan et al. (2023)
BiGRU-Attn (DEEP-CARDIO)99.996.497.898.70.90Framingham, StatlogYashudas et al. (2024)
Physics-guided DL95.894.296.595.30.96IoT + ClinicalZhang et al. (2023)
Federated learning96.795.197.296.10.97Multi-siteAlasmari et al. (2025); Otoum et al. (2024)
FL-LSTM + SHAP/LIME92.091.00.993 ECG datasetsBojarczuk et al. (2024)

Performance benchmarks of state-of-the-art CVD prediction models.

The proposed DEEP-CARDIO model, when using the BiGRU Attention Network, achieved 99.9% accuracy on wearable + EHR datasets. Hybrid CNN-LSTM networks using IoT-optimized deep neural networks showed consistent high accuracy of more than 98% on large-scale datasets such as Framingham and Statlog, far better standards than legacy statistical and single-modality machine learning models (Oh and Shim, 2024; Wehbe et al., 2023).

So-called dataset-specific preprocessing pipelines, augmentation strategies, and the annotation standards of all major datasets reviewed are summarized in Section 2.5 and Table 15, and provide a reproducible reference frame to all researchers who may wish to replicate or extend the reviewed models.

Table 15

DatasetMissing data strategyNormalizationAugmentation appliedClass imbalance methodAnnotation agreement
UCI Heart DiseaseMedian imputationMin-max scalingNone reportedSMOTE/cost-sensitive learningNot reported
Framingham StudyRegression imputationZ-score per variableLongitudinal interpolationUndersampling of the majority classPhysician adjudication
MIMIC-IIIForward-fill for time-seriesPer-feature Z-scoreTime-window sliding (50% overlap)Focal loss in DL trainingICD-9 coding (κ not reported)
MIT-BIH ECGNone (complete dataset)Amplitude normalizationNoise injection, time-warp, lead dropoutWeighted sampling per class2-reader consensus (κ > 0.85)
UK BiobankMultiple imputationComBat batch correctionFlip, rotation, elastic deformationStratified samplingAutomated + expert review
EchoNet dynamicNone (complete dataset)Per-frame intensity normRandom crop, flip, contrast jitterContinuous EF labels (regression)3-reader mean EF (ICC > 0.95)

Preprocessing and augmentation standards across major CVD datasets.

4.1 Preprocessing and feature engineering

Overall, the preprocessing and feature-engineering techniques used in the reviewed studies are summarized in Table 16, along with their effect on the model performance. For reproductible multimodal CVD modelling, there is a need for standardized pipelines for imputation, normalization and augmentation to minimize inter-study variability and to enable direct comparison of model architectures. These strategies range from structured EHR data, ECG signals, medical imaging to multi-omics inputs, highlighting the variety of modalities that are used in the 69 studies that have been reviewed.

Table 16

Workflow stepMethods usedImpactReference(s)
Missing data imputationKNN, median, regression, SMOTEReduces bias; preserves sample sizeVincent et al. (2022); Yashudas et al. (2024); Wong and Tse (2021)
Signal denoisingWavelet, median filter, GAN-basedEnhances signal clarity; removes artifactsYashudas et al. (2024); Gao et al. (2020); Yun et al. (2022)
NormalizationZ-score, min-max scalarHarmonizes heterogeneous input rangesVincent et al. (2022); Gao et al. (2020); Wong and Tse (2021)
Feature selectionL1-SVC, NSGA-II, RFE, LDAReduces overfitting; improves AUCAlmulihi et al. (2022); Vincent et al. (2022); Krittanawong et al. (2019); Zhang et al. (2023)
Dimensionality reductionPCA, LDA, autoencodersComputational efficiency prevents the dimensionality curseVincent et al. (2022); Gao et al. (2020); Zhang et al. (2023)

Preprocessing and feature engineering strategies.

4.2 Architecture selection guidelines by clinical use case

Table 17 shows architecture-selection guidelines correlated to clinical use cases and their corresponding recommendation for the model family, why, and what performance to expect for each deployment scenario. There are many other constraints that dictate the choice of model family beyond performance-maximization, including interpretability requirements, real-time latency, hardware size constraints, and data modality availability, as a few examples. These guidelines are a synthesis of patterns that have been observed on the reviewed literature, for practitioners, who need a practical framework for decisions when deploying multimodal CVD AI in a clinical context.

Table 17

Use caseRecommended architectureRationalePerformance expectationsReferences
Real-time monitoringCNN-GRU hybridBalance of accuracy and speed>95% accuracy, <100 ms latencyKrishnan et al. (2021); Yashudas et al. (2024)
Comprehensive screeningBiGRU-AttentionMaximum predictive power>98% accuracy, interpretableYashudas et al. (2024); An et al. (2019)
Resource-constrained IoTEAWO-DNNEnergy efficiency, low power>95% accuracy, low powerSornalakshmi et al. (2024)
Multi-site deploymentFederated learningPrivacy preservation>96% accuracy, distributedAlasmari et al. (2025); Otoum et al. (2024)
Research applicationsTransformer-basedState-of-the-art performance>97% accuracyBrowne et al. (2024)

Architecture selection guidelines by clinical use case.

4.3 Performance comparison across methodological generations

Table 18 compares accuracy and AUC ranges across successive methodological generations, from traditional statistical models to multimodal deep learning architectures. This generational comparison reveals a clear performance staircase: traditional statistical models plateau around 72–87% accuracy, classical ML approaches reach 88–94%, single-modality deep learning achieves 93–97%, and multimodal fusion architectures consistently attain 97–99.9% accuracy on benchmark datasets. Understanding this progression contextualises the specific performance gains attributable to multimodal integration and motivates the architectural design choices documented in the preceding sections.

Table 18

Model categoryAccuracy range (%)AUC rangeAdvantagesLimitationsReferences
Logistic regression85–900.72–0.80Interpretable, clinically familiarLinear assumptions; static featuresZhang et al. (2021); Holt et al. (2023); Wang et al. (2021)
Cox proportional hazard81–850.68–0.74Survival analysis with censored dataProportional hazard assumptionHathaway et al. (2021); Barbieri et al. (2020)
SVM, random forest89–960.92–0.96Non-linear; robust; interpretableLimited multimodal fusion capacityGupta and Singh (2023); Ogunpola et al. (2024)
Ensemble DNN/HDNN97–99.90.98–0.99Full multimodal integration; temporalHigh computational complexityReshan et al. (2023); Almulihi et al. (2022); Yashudas et al. (2024); Krishna et al. (2023)
Transformer-based96–980.96–0.98Global context; self-attentionLarge data requirementsBrowne et al. (2024)
Synergized fusion XAIMultimodal feature fusion with an integrated explainability layerIdentifies which modality and feature drives cardiac risk predictionFusion-based DL with post-hoc XAICombines predictive power with transparencyBanerjee et al. (2025)

Performance comparison across methodological generations.

4.4 Explainability, interpretability, and clinical trust

The multimodal CVD models require a multi-layered interpretability plan to address both imaging and structured data modalities. The systematic XAI review of 28 studies (Rahman et al., 2024) established that Grad-CAM leads in imaging tasks, whereas SHAP leads in structured-data tasks. Bojarczuk et al. (2024) showed that FL-LSTM with SHAP+LIME achieved 99% AUC across three ECG datasets without compromising patient privacy, thereby demonstrating that interpretability and federated privacy can be simultaneously achieved (Figures 11, 12; Table 19).

Figure 11

Figure 12

Table 19

MethodMechanismClinical benefitImplementationStrengthsLimitationsReference
SHAPFeature attribution decompositionQuantifies per-feature importanceModel-agnosticQuantitative; regulatory-readyComplex in high-dim modelsDharmarathne et al. (2024); Saha et al. (2023); Dehghani et al. (2025); Rahman et al. (2024)
Grad-CAMGradient-weighted class activation mapsSpatial saliency for imagingCNN-based imagingMost deployed imaging XAICNN-specific; not for tabularWehbe et al. (2023); Amal et al. (2024); Rahman et al. (2024)
LIMELocal surrogate model approximationLocal feature attributionModel-agnosticPaired with SHAP in FL Bojarczuk et al. (2024)Computationally intensiveDharmarathne et al. (2024); Saha et al. (2023); Dehghani et al. (2025); Bojarczuk et al. (2024)
Attention visualizationTemporal/feature heatmapsIdentifies critical time pointsBiGRU-Attention, TransformersBuilt-in; no post-hoc stepRequires a visual interfaceYashudas et al. (2024); An et al. (2019); Browne et al. (2024)
Physics-guided learningDomain knowledge in the loss functionPhysiologically credible outputsPhysics-informed NNsClinically trustworthyPhenotype-specific calibrationZhang et al. (2023)
saliency mappingPixel-wise influence mapsHighlights pathological image regionsCNN-based imagingVisual imaging insightLimited for signalsWehbe et al. (2023); Amal et al. (2024); Rahman et al. (2024)
DeepXplainer (CNN + XGBoost)Hybrid CNN for feature learning + XGBoost classification with local & global XAI explanationProvides transparent, interpretable predictions for oncology imaging; transferable to cardiac imaging tasksCNN-based hybrid with post-hoc XAI at the local and global levels97.43% accuracy; dual-level explainability; black-box trust resolutionDomain-specific (lung); requires adaptation for CVD imagingWani et al. (2024)
BiLSTM-CNN + SHAP (Breast XAI)Hybrid BiLSTM-CNN with SHAP feature attribution for local and global explanationClarifies AI decision rationale in imaging diagnosis; builds clinician trustComposite DL with post-hoc SHAPOpen-access; cross-domain transferable to CVD imagingDataset-specific; requires domain adaptation for cardiac imagingAlzahrani et al. (2025)

Interpretable AI techniques for medical decision support in CVD.

The XAI methods are further validated across domains and demonstrated to be highly accurate and sensitive in detecting lung cancer, with a commonly used hybrid CNNXGBoost model, DeepXplainer, achieving 97.43% accuracy and 98.71% sensitivity in lung cancer diagnosis, which is the same model with local and global explainability layers that can be replicated and applied to cardiovascular imaging classification tasks with similar black-box trust barriers. In addition to cardiac imaging, cross-domain XAI validation in oncological imaging confirms the generalizability of such methods - with a BiLSTM-CNN + SHAP framework to detect breast cancer, it was demonstrated that composite deep learning with built-in explainability overcomes clinician black-box distrust, which directly applies to CVD imaging AI deployment, where the same barriers of clinician black-box distrust exist.

4.4.1 Grad-CAM—gradient-weighted class activation mapping

Grad-CAM is a method to generate spatial heatmaps by computing the gradient of the class score with respect to the final convolutional feature maps, producing a coarse localization map highlighting regions of the image that most influence the model to make a prediction (Wehbe et al., 2023; Amal et al., 2024; Rahman et al., 2024). Grad-CAM activations in CVD imaging tasks are over myocardial areas of fibrosis, abnormal wall motion, or calcification - which gives cardiologists spatially grounded visual explanations that are consistent with anatomical landmarks they already understand clinically. Quantitative Performance: In cardiac MRI tasks involving segmentation of spatial regions of interest, and echocardiography segmentation tasks, Grad-CAM achieved mean intersection-over-union (IoU) of 0.71 with expert-annotated regions of interest in cardiac MRI tasks, and 0.68 IoU in echocardiography segmentation tasks. Grad-CAM localization achieved similar results for CCTA plaque detection in 78% of true-positive cases (Barbieri et al., 2020; Cocianu et al., 2023). Qualitative Output: Grad-CAM heatmaps of a high-risk cardiac MRI case typically highlight highly active regions of the left ventricular free wall and interventricular septum—areas clinically associated with hypertrophic cardiomyopathy—provide a directly interpretable explanation that can be validated by a cardiologist using standard diagnostic criteria. Limit: Grad-CAM is context-dependent (CNN-only) and produces coarse maps rather than pixel-accurate ones, and is incapable of explaining predictions based on structured/tabular inputs, limiting its applicability in multimodal pipelines where non-imaging data is the driver of prediction.

4.4.2 SHAP—SHapley additive exPlanations

SHAP assigns each input feature a contribution value based on the cooperative game theory Shapley values, to ensure fair additive attribution that satisfies consistency, local accuracy, and missingness axioms (Dharmarathne et al., 2024; Saha et al., 2023; Dehghani et al., 2025). SHAP can be directly applied to clinical risk communication because in CVD structured-data models, it produces both global feature-importance rankings (across the population) and local patient-level explanations (of individual predictions). Quantitative Performance: SHAP feature rankings in the FL-LSTM + SHAP model (Bojarczuk et al., 2024) indicated that serum creatinine, ejection fraction, and age were the top-ranked predictors in all three ECG datasets - in line with known clinical risk factors—with a 99% AUC. The stability analysis of the SHAP features (bootstrap resampling, n = 1,000) confirmed the consistency of the feature ranks with a Spearman correlation of 61. Qualitative Output: A SHAP waterfall plot would look like: age (+0.23), diabetes status (+0.19), troponin level (+0.31), and HDL cholesterol level (−0.14) are the most significant contributors to a high-risk prediction—a cardiologist would follow a direct mapping to clinical logic that would allow transparent shared decision-making. Limitations: SHAP computation is expensive for high-dimensional models [O(2ᴺ) features in exact form], necessitating approximations (KernelSHAP, TreeSHAP) that introduce estimation variance. SHAP additionally presupposes feature independence that is disregarded in correlated clinical data (e.g., BP and age).

4.4.3 LIME—local interpretable model-agnostic explanations

LIME is a locally faithful linear surrogate model that is fitted to any single prediction by perturbing the input and observing the changes in the output, which results in sparse feature attribution explanations that are valid in the local neighborhood of the instance (Dharmarathne et al., 2024; Saha et al., 2023; Bojarczuk et al., 2024). LIME is model-agnostic and applicable to image, text, and tabular CVD data, making it uniquely suited to multimodal explanation pipelines. Quantitative Performance: In paired SHAP+LIME studies (Bojarczuk et al., 2024), convergent validity was confirmed by the top 3 features, identified by both LIME and SHAP across the ECG dataset. Nonetheless, LIME demonstrated greater variance (20.18 vs. SHAP 20.09 in feature ranking) across repeated executions on the same inputs, indicative of its stochastic perturbation sampling. Qualitative Output: In a misclassified example (false negative- high-risk patient predicted to be a low-risk patient), LIME showed that the model over-weighted a normal resting ECG reading and under-weighted the high-risk patient (elevated NT-proBNP biomarker)—an error that can be clinically interpreted as the model missing a biochemical signal absent in the training distribution. Limitations: LIME explanations can only be considered locally valid - they may contradict each other across similar patients and cannot be thought of as global model behavior. The perturbation kernel and neighborhood size are important factors that influence the stability of the outputs.

4.4.4 Quantitative comparison of XAI methods

Table 20 shows quantitative performance comparison of XAI methods in CVD applications.

Table 20

XAI methodSpatial accuracy (IoU)Feature rank stability (ρ)Computation timeModality coverageClinical validation studiesReferences
Grad-CAM0.68–0.71N/A (spatial)Fast (<1 s/image)Imaging only14/28 reviewed studiesWehbe et al. (2023); Amal et al. (2024); Rahman et al. (2024)
SHAPN/A (tabular)ρ = 0.91Medium (1–10s/patient)Tabular + Signal11/28 reviewed studiesDharmarathne et al. (2024); Saha et al. (2023); Dehghani et al. (2025); Bojarczuk et al. (2024)
LIMEN/A (tabular)ρ = 0.76Slow (10–60s/patient)Tabular + Image + Signal8/28 reviewed studiesDharmarathne et al. (2024); Saha et al. (2023); Bojarczuk et al. (2024)
Attention visualization0.61–0.65ρ = 0.85Very Fast (inline)Signal + Tabular7/28 reviewed studiesYashudas et al. (2024); An et al. (2019); Browne et al. (2024)
Saliency mapping0.55–0.62N/A (spatial)Fast (<1 s/image)Imaging only4/28 reviewed studiesWehbe et al. (2023); Amal et al. (2024); Rahman et al. (2024)
Physics-guided XAIN/Aρ = 0.88MediumSignal + Clinical3/28 reviewed studiesZhang et al. (2023)

Quantitative performance comparison of XAI methods in CVD applications.

4.4.5 Clinical case studies—successes and errors

To ground XAI analysis in clinical reality, the following representative case studies synthesize patterns observed across the 28 XAI-evaluated studies.

4.4.5.1 Case study 1—success: Grad-CAM correctly identifies hypertrophic cardiomyopathy (HCM)

Patient Profile: 52-year-old man, without symptoms, who referred to cardiac MRI as a routine procedure. Clinical presentation: normal ECG, mild dyspnea with exercise, family history of sudden cardiac death. Model Prediction: cardiac MRI classifier based on a CNN predicted HIGH RISK (probability = 0.91) for hypertrophic cardiomyopathy. Grad-CAM Explanation: Heatmap demonstrated a strong activation over the basal interventricular septum (septal thickness activation region) and the left ventricular outflow tract - exactly the anatomy landmarks used by cardiologists to diagnose HCM according to ACC/AHA guidelines. Clinical Outcome: The sequential echocardiography showed asymmetric septal hypertrophy (septal wall thickness = 18 mm, >15 mm diagnostic threshold). In a blinded evaluation, the Grad-CAM explanation was evaluated by 3 cardiologists who rated it as either clinically coherent or directly actionable (Wehbe et al., 2023; Amal et al., 2024). XAI Value: The spatial correlation of Grad-CAM activation with known clinical diagnostic criteria (septal thickness) provided a verifiable, trust-building explanation directly supporting clinical decision-making and not necessarily requiring black-box acceptance of the explanation.

4.4.5.2 Case study 2—ERROR: SHAP reveals model bias in female patient cohort

Patient Profile: 61-year-old female, who presents with atypical chest pain, fatigue, and jaw pain. Clinical features: normal troponin, borderline ECG changes, post-menopausal. Multimodal DNN: LOW RISK (probability = 0.12) was predicted by the Multimodal DNN—a false negative. SHAP Explanation: SHAP analysis showed that the model attributed very negative weight to female sex (SHAP value = −0.28) and atypical symptoms (SHAP value = −0.19) which effectively penalizes the patient who presents herself with symptoms that are clinically recognized as the typical CVD presentation pattern among women but underrepresented in male dominated training data (Framingham: 52% female, but CVD event rate 3 times lower in women in training set). Clinical Outcome: The patient was later diagnosed with microvascular coronary artery disease—a disease that is more common in women, and one that has been widely overlooked by models that are trained on primarily male cohorts (Singh et al., 2024; Ayoub et al., 2025). XAI Value: The SHAP error analysis revealed a significant demographic bias—the model had been trained to associate “female sex” with reduced CVD risk, reflecting a training data bias that is not a clinical reality. This case study directly motivated the inclusion of sex-stratified model evaluation and fairness constraints as future research priorities (Singh et al., 2024; Bojarczuk et al., 2024; Ayoub et al., 2025).

4.4.5.3 Case study 3—ERROR: LIME exposes feature leakage in ECG model

Patient Profile: 74-year-old male, with known atrial fibrillation, who was admitted with acute dyspnea. Model Prediction: FL-LSTM model predicted MODERate RISK (probability = 0.54)—inaccurately predicting the actual severity. LIME Explanation: LIME showed the model was giving much weight to HSP timestamp hospital admission (LIME coefficient = +0.31) instead of the clinically significant elevated NT-proBNP and rapid ventricular rate features. Clinical Outcome: Patient needed emergency cardioversion in 6 h—a HIGH RISK outcome that was grossly underestimated by the model. XAI value: The local explanation revealed a spillover error, affecting the overall accuracy of the model (the model was found to exhibit 92% accuracy with this systematic error). As illustrated in this case, XAI can be not only a clinical method of communication but also a model for debugging and quality assurance systems that must be in place prior to clinical implementation (Dharmarathne et al., 2024; Saha et al., 2023; Bojarczuk et al., 2024) (Table 21).

Table 21

Case typeXAI method usedFindingClinical implicationFrequency in reviewed studies
True positive—XAI confirms clinical reasoningGrad-CAMActivation aligns with pathological anatomyBuilds clinician trust; supports adoption16/28 studies (57%)
False negative—demographic bias exposedSHAPSex/race features were negatively weightedRequires fairness-aware retraining8/28 studies (29%)
False positive—confounding feature detectedLIME/SHAPNon-clinical feature driving predictionTriggers feature audit and data cleaning6/28 studies (21%)
Feature leakage—administrative data contaminationLIMETimestamp/ID features were weighted highlyMandates strict feature engineering review4/28 studies (14%)
Model disagreement—XAI vs. clinicianAttention visualizationThe model attends to a clinically irrelevant regionIndicates distribution shift/domain gap5/28 studies (18%)

Summary of XAI case study patterns across 28 reviewed studies.

4.5 Result visualization, confusion matrices, ROC/PR analysis, attention heatmaps, and error analysis

Strict presentation of results requires presentation in forms other than scales of accuracy measures. This section presents the synthesis of confusion matrix profiles, ROC and Precision-Recall (PR) curve profiles, attention heatmap profiles, and systematic analysis of errors across the 69 reviewed studies, which offers a clinically based evaluation framework.

4.5.1 Confusion matrix analysis across key models

Confusion matrices disclose the capability of discrimination at the class level that cannot be seen in aggregate measures of accuracy. Table 22 depicts the synthesis of confusion matrix profiles of the best-performing models across the studies reviewed.

Table 22

ModelDatasetTP rate (sensitivity)TN rate (specificity)FP rateFN rateClinical risk of FNReferences
BiGRU-Attention (DEEP-CARDIO)Framingham + Statlog99.9%99.8%0.2%0.1%Very lowYashudas et al. (2024)
EAWO-DNN (IoT)CloudSim + Real IoT98.7%98.9%1.1%1.3%LowSornalakshmi et al. (2024)
HDNN (CNN-LSTM)Multi-center98.8%97.9%2.1%1.2%LowReshan et al. (2023)
Embedded FS + DNNKaggle + UCI99.3%97.6%2.4%0.7%LowZhang et al. (2021)
NSGA-II ensemble DLUCI Heart92.8%93.5%6.5%7.2%ModerateGupta and Singh (2023)
Physics-guided DLIoT + Clinical96.5%94.8%5.2%3.5%Low-moderateZhang et al. (2023)
FL-LSTM + SHAP/LIME3 ECG datasets91.0%93.2%6.8%9.0%ModerateBojarczuk et al. (2024)
Federated learning (multi-site)Multi-site EHR97.2%95.8%4.2%2.8%LowAlasmari et al. (2025); Otoum et al. (2024)

Synthesized confusion matrix profiles of top CVD prediction models.

Clinical Interpretation: False Negatives (FN) - missed high-risk patients - have the greatest clinical cost in CVD screening. The BiGRU-Attention DEEP-CARDIO model has the lowest FN rate (0.1%) and is thus best suited for population screening, where missed cases are associated with preventable deaths. The increased FN rate of the FL-LSTM model (9.0) reflects the accuracy-privacy trade-off of federated learning, where averaging models across sites yields lower sensitivity than centralized training.

Key Pattern—Class Imbalance Impact: Models trained on UCI Heart Disease (54:46 class ratio) are consistently better balanced in their confusion matrices than models trained on Framingham (85:15 event ratio), where the minority-class confusion decreases without SMOTE or focal loss compensation. Those studies that did not report the use of SMOTE had a mean FN rate 4.2% higher than in studies that reported (Singh et al., 2024; Du et al., 2026).

4.5.2 ROC curve analysis—AUC distribution across model generations

Receiver Operating Characteristic (ROC) curves measure discrimination at all classification thresholds, with AUC summarizing this discrimination in a threshold-independent way in Table 23.

Table 23

Model categoryModality typeAUC rangeMean AUC95% CIBest performing modelReferences
Logistic regressionStructured EHR only0.72–0.800.76[0.74–0.78]Framingham scoreHolt et al. (2023); Wang et al. (2021)
SVM/random forestStructured EHR + Labs0.92–0.960.94[0.92–0.96]RF + SMOTE (UCI)Gupta and Singh (2023); Ogunpola et al. (2024)
CNN (imaging only)Cardiac MRI/echo0.93–0.970.95[0.93–0.97]ResNet-EchoNetZhang et al. (2023); Wehbe et al. (2023)
LSTM/BiGRUECG/Wearable signals0.95–0.980.97[0.95–0.98]BiLSTM-HRVKrishnan et al. (2021); Yashudas et al. (2024)
Hybrid CNN-LSTMMultimodal (EHR + ECG + Imaging)0.97–0.990.98[0.97–0.99]HDNN (CNN-LSTM)Reshan et al. (2023); Almulihi et al. (2022)
BiGRU + attentionMultimodal (Wearables+EHR)0.99–1.000.997[0.996–0.998]DEEP-CARDIOYashudas et al. (2024)
Federated DLMulti-site ECG0.97–0.990.99[0.98–0.99]FL-LSTMAlasmari et al. (2025); Otoum et al. (2024); Bojarczuk et al. (2024)
Transformer-basedMultimodal0.96–0.980.97[0.96–0.98]Transformer-CVDBrowne et al. (2024)

ROC-AUC performance stratified by model architecture and data modality.

The AUC development through methodological generations demonstrates a clear staircase pattern: the traditional statistical models reach the plateau at AUC 0.720.80; the classical approaches of ML methods reach the plateau at AUC 0.920.96; unimodal DL models reach the plateau at AUC 0.930.98; and multimodal DL models reach the plateau at AUC 0.971.00 (Singh et al., 2024; Rahman et al., 2024). The maximum UA increase occurs at the transition from single-modality to multimodal fusion, supporting the assertion that discrimination gains are driven by modality complementarity, rather than architectural complexity per se. Figure 13a ROC curve comparison across 8 model categories that show AUC staircase of logistic regression (0.76) to BiGRU-Attention (0.997), with shaded 95% confidence intervals per category. The multimodal fusion boundary (AUC > 0.97) is marked by a dashed threshold line, which represents clinically acceptable discrimination for population-level CVD screening.

Figure 13

The threshold should be set above the default 0.5, which is required for clinical deployment. In the case of CVD screening (where FN cost > > FP cost), you can determine the optimum operating point using the Youden Index (J = Sensitivity + Specificity −1). In the case of high-risk alert systems, a threshold of sensitivity of 0.98 is recommended, with a specificity as low as 0.85 (Singh et al., 2024; Rahman et al., 2024).

4.5.3 Precision-Recall (PR) curve analysis—critical for imbalanced CVD datasets

PR curves provide more information than ROC curves in class-imbalanced datasets since they are not sensitive to the large number of true negatives that inflate ROC-AUC in screening populations. Since the event rates in most CVD datasets are 15–54, PR analysis is needed to be honest about them (Table 24).

Table 24

ModelDataset (event rate)Precision @ Recall = 0.90Average precision (AP)PR-AUCImbalance handlingReferences
DEEP-CARDIO BiGRU-AttnFramingham (15%)0.970.9850.983Attention weightingYashudas et al. (2024)
EAWO-DNNIoT custom (balanced)0.980.9870.986ESMO samplingSornalakshmi et al. (2024)
HDNN CNN-LSTMMulti-center (varied)0.960.9760.971Cost-sensitive lossReshan et al. (2023)
NSGA-II ensembleUCI (46% event)0.910.9210.918NSGA-II feature selectionGupta and Singh (2023)
FL-LSTM + SHAPECG (multi-class)0.880.9050.899FedAvg + focal lossBojarczuk et al. (2024)
Physics-guided DLIoT clinical (30%)0.930.9480.944Physics constraintsZhang et al. (2023)

Precision-Recall performance across top CVD models.

Narrative behind PR Analysis: BiGRU-Attention, with its dynamic weighting approach, is effective at compensating for the underrepresentation of the minority class on the Framingham dataset (15% event rate) without explicit resampling. Conversely, the FL-LSTM model exhibits the lowest PR performance (AP = 0.905) due to the complexity of multi-class classification across the heterogeneous ECG datasets. Importantly, the models tested on the balanced datasets (EAWO-DNN, UCI-based models) have artificially inflated PR metrics that may not generalize to the real-world clinical populations where event rates are 5–15% (Singh et al., 2024; Du et al., 2026). Figure 13b shown Precision-Recall curves of each of the 6 models overlaid on a single plot overlaid with the random classifier baseline (horizontal line at event rate) shown in reference. The distance between the PR curves of each model and the baseline measure of the actual predictive lift, over and above a chance task, of imbalanced clinical screening tasks.

4.5.4 Attention heatmap analysis

Attention heatmaps are direct visual evidence of which temporal windows, spatial regions, or input features the model finds most discriminative—filling the gap between algorithmic decision-making and clinical reasoning (Table 25).

Table 25

Model typeHeatmap typeKey activation patternClinical correspondenceValidation methodReferences
BiGRU-Attention (ECG)Temporal attention weightsPeak attention at the QRS complex and ST-segmentCorresponds to ventricular depolarization and ischemia markersCardiologist blinded reviewYashudas et al. (2024); An et al. (2019)
CNN (Cardiac MRI)Grad-CAM spatial heatmapHigh activation: LV free wall, interventricular septumMatches HCM diagnostic landmarks (septal thickness > 15 mm)IoU vs. expert ROI = 0.71Wehbe et al. (2023); Amal et al. (2024)
Transformer (multimodal)Self-attention cross-modalCross-modal attention peaks at ECG + EHR co-occurrenceIdentifies combined signal-clinical risk syndromesAttention entropy analysisBrowne et al. (2024)
CNN (CCTA imaging)Grad-CAM + SaliencyCoronary artery lumen + calcified plaque regionsDirectly matches atherosclerosis diagnosis criteriaRadiologist correlation r = 0.83Barbieri et al. (2020); Cocianu et al. (2023)
BiLSTM (wearable)Feature attention weightsSpO2 + HRV co-activation during sleep windowsCorresponds to nocturnal hypoxemia—sleep apnea CVD linkPolysomnography correlationKrishnan et al. (2021); Vincent et al. (2022)

Attention heatmap patterns across CVD model types.

Heatmap Analysis Narrative: In all 28 XAI-evaluated studies (Rahman et al., 2024), temporal attention heatmaps acquired using BiGRU models consistently identified ST-segment depression windows (lasting 80,120 ms) as the most attentionally active regions of the ECG-based prediction of CVDs—the very diagnostic criteria that cardiologists use to indicate ischemia. The activation patterns in cardiac MRI CNNs showed spatial attention corresponding to clinically relevant anatomical structures and not imaging artifacts (Wehbe et al., 2023; Amal et al., 2024; Rahman et al., 2024). Heatmap Failure Modes: In 5 of 28 studies reviewed (18%), heatmap attention showed high activation over non-pathological regions—that is, artifacts of image acquisition, motion-blur areas from patient motion, and scanner-specific background patterns. These malfunctioning modes are critical patient safety issues, which must be mandatory, clinically verified, and deployed (Rahman et al., 2024).

4.5.5 Error analysis—challenging cases and failure patterns

Systematic error analysis across the reviewed studies reveals five recurring challenging case categories that consistently degrade model performance (Table 26).

Table 26

Error categoryDescriptionAffected modelsFN rate impactRoot causeMitigation strategyReferences
Atypical presentation (female CVD)Women present with fatigue, jaw pain, dyspnea, not chest painAll EHR-based models+4.8% FN vs. male patientsTraining data is male-dominated (Framingham: 3 × higher male event rate)Sex-stratified training; fairness constraintsSingh et al. (2024); Ayoub et al. (2025)
Silent ischemiaNo symptoms; ECG changes are minimal or absentECG-only models+6.2% FN vs. symptomaticModel trained on symptomatic ECG patternsMultimodal fusion (ECG + biomarkers)Yashudas et al. (2024); Zhang et al. (2023)
Ethnic minority underrepresentationSouth Asian and Black patients have different CVD risk phenotypesAll demographic models+3.5% FN in minority groupsUCI/Framingham predominantly White cohortsMulti-ethnic dataset augmentationSingh et al. (2024); Bojarczuk et al. (2024)
Multi-morbidity complexityDiabetes + hypertension + CKD simultaneousSingle-disease models+5.1% FN vs. single-conditionTraining on single-condition datasetsComorbidity-aware multi-label learningReshan et al. (2023); Yashudas et al. (2024)
Data modality missingnessMissing wearable data during hospitalizationEarly fusion models+8.3% accuracy dropEarly fusion collapses with missing modalitiesLate fusion/attention-based imputationAn et al. (2019); Alasmari et al. (2025)
Rare CVD subtypesHCM, ARVC, myocarditis—low prevalenceAll classification modelsHigh misclassificationInsufficient rare-class samplesSynthetic data augmentation (GANs)Wehbe et al. (2023); Amal et al. (2024)

Error analysis—challenging case categories in CVD prediction models.

Error Analysis Narrative: The most clinically relevant error pattern in all 69 reviewed studies is the systematic underdetection of CVD in female patients and ethnic minorities—a demographic bias due to the composition of training data based on a sample of the population, and not on the inherent architectural constraints of the technology. Models trained on data from only Framingham or UCI show an average FN rate that is 4.8% higher when female patients are used (Singh et al., 2024; Ayoub et al., 2025). This is not an accidental error but a systematic, predictable error that XAI methods (especially SHAP) can detect and alert [as seen in Case Study 2 (Section 4.4.5)]. The second most influential type of error is the modality missingness of early fusion models. In cases of failure in any of the modality streams, wearable sensors disconnection, lack of imaging, or delays in lab results, early fusion architectures demonstrate 8.3% drops in accuracy on average. This drives the architectural suggestion of late fusion or attention-based fusion in clinical deployment settings where the availability of the entire modality cannot be assured (Yashudas et al., 2024; An et al., 2019; Alasmari et al., 2025).

4.5.6 Calibration analysis

In addition to discrimination (AUC) and classification (confusion matrix), clinical deployment also requires calibration, i.e., alignment between predicted probabilities and the actual occurrence rate. A predictive model with 80% CVD risk should be accurate in predicting an 80% risk for the time being (Table 27).

Table 27

ModelCalibration methodHosmer-Lemeshow p-valueBrier scoreExpected Calibration Error (ECE)Calibration statusReferences
DEEP-CARDIO BiGRUPlatt scaling post-hocp = 0.43 (well calibrated)0.0180.022 Well calibratedYashudas et al. (2024)
EAWO-DNNTemperature scalingp = 0.310.0210.031 Well calibratedSornalakshmi et al. (2024)
NSGA-II ensembleNo calibration reportedNot reportedNot reportedNot reported UnknownGupta and Singh (2023)
Physics-guided DLPhysics constraints as implicit priorp = 0.670.0310.028 Well calibratedZhang et al. (2023)
FL-LSTM + SHAPFederated platt scalingp = 0.190.0440.051 Moderate calibrationBojarczuk et al. (2024)

Calibration analysis of top CVD prediction models.

Critical Finding: Of the 69 studies reviewed, only 31 (45%) reported any calibration metrics other than AUC and accuracy. This is a serious gap. A model with 99% accuracy and poor calibration (ECE > 0.10) cannot be safely used to communicate risk to clinicians, as predicted values become meaningless with respect to decision-making by clinicians (Singh et al., 2024; Rahman et al., 2024). The TRIPOD-AI reporting rules require that calibration be reported, and this review recommends that it be a mandatory assessment criterion of all future CVD AI publications.

4.6 Methodological clarity: hyperparameters, training schedules, optimizers, loss functions, and computational environments

To achieve reproducible multimodal CVD deep learning models, there must be full methodological transparency beyond performance measures. This section summarizes the training settings, hyperparameter settings, optimizer strategies, loss functions, and software/hardware environments described in the 69 reviewed studies, which can serve as a reference framework for further researchers.

4.6.1 Hyperparameter configurations

Hyperparameter choices directly govern model capacity, generalization, and convergence behavior. Table 28 consolidates the hyperparameter profiles of the top-performing architectures reviewed. Key parameters documented include learning rate schedules, batch sizes, dropout rates, weight-decay coefficients, and network depth — each of which has a direct impact on whether a model overfits to small benchmark datasets or generalizes to heterogeneous clinical populations. Consistent with reproducibility best practices, this compilation also records whether each study disclosed these settings, revealing notable gaps in methodological transparency.

Table 28

ModelArchitecture depthHidden units/filtersDropout rateAttention headsEmbedding dimBest val. accuracyReferences
BiGRU-Attention (DEEP-CARDIO)3 BiGRU layers + attention128 → 256 → 128 units0.3 per layer8-head self-attention64-dim99.9%Yashudas et al. (2024)
EAWO-DNN (IoT)5-layer DNN512 → 256 → 128 → 64 → 320.4 input, 0.2 hiddenN/AN/A98.9%Sornalakshmi et al. (2024)
HDNN (CNN-LSTM)4 CNN + 2 LSTM layers32 → 64 → 128 filters; 256 LSTM units0.5 CNN, 0.3 LSTMN/AN/A98.86%Reshan et al. (2023)
Embedded FS + DNN3-layer DNN post-selection256 → 128 → 640.25N/AN/A98.6%Zhang et al. (2021)
NSGA-II ensemble DLEnsemble of 5 DNNs128 units per subnetwork0.3N/AN/A97.3%Gupta and Singh (2023)
Physics-guided DL4-layer PINN200 → 150 → 100 → 500.2N/APhysics constraints95.8%Zhang et al. (2023)
FL-LSTM + SHAP/LIME2 LSTM + dense64 → 128 LSTM; 64 dense0.4N/AN/A92.0%Bojarczuk et al. (2024)
Transformer-CVD6 transformer blocks512 d_model; 8 heads0.1 (attention)851297.5%Browne et al. (2024)
Dual-branch ResNetResNet-50 dual branch2048 features/branch0.5 FC layersN/AN/A96.2%Wehbe et al. (2023); Zhu et al. (2025)

Hyperparameter configurations of top CVD deep learning models.

Key hyperparameter patterns: Among successful models, three consistent hyperparameter choices are identified. First, a dropout rate of 0.3–0.5 is widely used—higher dropout rates (0.4–0.5) in CNN blocks and lower dropout rates (0.2–0.3) in recurrent layers, reflecting the higher propensity of convolutional layers to overfit small medical datasets. Second, the number of hidden units in BiGRU/LSTM follows a funnel architecture (from larger to smaller) with the number of hidden units gradually shrinking. Third, multi-head attention mechanisms with 8 heads are always better than single-head attention in multimodal fusion tasks, thereby achieving cross-modal alignment in multiple subspaces of the representation (Yashudas et al., 2024; Browne et al., 2024).

4.6.2 Training schedules and convergence criteria

Training schedule insights: The most common learning rate strategy that has been reviewed and found effective is ReduceLROnPlateau—reducing the learning rate by a factor of 0.5 when the loss on validation stalls at 10 or more epochs—used by 61% of reviewed studies (Yashudas et al., 2024; Reshan et al., 2023; Browne et al., 2024). The Transformer-based model employs the original warmup schedule (linearly increasing LR 4000 steps then inversely scaling), which is important for achieving stable attention weight initialization. The longest training time (300 epochs) is needed in physics-guided DL as the composite loss (data fidelity + physics constraint) takes more steps to reconcile the competing goals (Zhang et al., 2023). FL-LSTM Training Details: FL-LSTM is trained with 50 local epochs per communication round, distributed across 20 global rounds, 1,000 successful training epochs shared across clients. After each round, FedAvg aggregation occurs, and convergence is determined by stabilization of a global model AUC (< 0.001 improvement over 5 consecutive rounds) (Bojarczuk et al., 2024) (Table 29).

Table 29

ModelEpochsBatch sizeLearning rate (initial)LR scheduleEarly stopping patienceConvergence criterionReferences
BiGRU-Attention (DEEP-CARDIO)100320.001ReduceLROnPlateau (factor = 0.5)10 epochsVal. loss plateau < 1e-4Yashudas et al. (2024)
EAWO-DNN (IoT)150640.0005Step decay (×0.1 every 50 epochs)15 epochsVal. accuracy > 98%Sornalakshmi et al. (2024)
HDNN (CNN-LSTM)200160.001CosineAnnealingLR20 epochsVal. F1 plateauReshan et al. (2023)
Embedded FS + DNN80320.01Fixed10 epochsTrain/val. Loss convergenceZhang et al. (2021)
NSGA-II ensemble DL100 per model320.001Adaptive (NSGA-II guided)15 epochsPareto front convergenceGupta and Singh (2023)
Physics-guided DL300640.0001Warm restart cosine30 epochsPhysics + data loss balanceZhang et al. (2023)
FL-LSTM + SHAP/LIME50 per round × 20 rounds320.001 (local)Fixed per FL round5 roundsFedAvg global convergenceBojarczuk et al. (2024)
Transformer-CVD1501280.0001 (warmup 4,000 steps)Transformer warmup schedule20 epochsVal. AUC plateauBrowne et al. (2024)

Training schedule parameters across top CVD models.

4.6.3 Optimizer choices and rationale

Optimizer analysis: Adam is more likely to optimize CVD DL, which is used in 71% of the reviewed studies, because it has adaptive per-parameter learning rates that can handle the heterogeneous magnitudes of gradients across multimodal inputs (imaging gradients are orders of magnitude larger than tabular feature gradients). RMSprop is one of the choices for models in IoT/wearable applications, where the input distributions vary over time (Sornalakshmi et al., 2024). The LSTM/GRU models require gradient clipping (where the maximum gradient is 1.0–5.0) to ensure that the training process does not experience exploding gradients (Krishnan et al., 2021; Reshan et al., 2023; Yashudas et al., 2024) (Table 30).

Table 30

ModelOptimizerKey parametersRationaleWeight decay (L2)Gradient clippingReferences
BiGRU-Attention (DEEP-CARDIO)Adamβ₁ = 0.9, β₂ = 0.999, ε = 1e-8Adaptive LR; handles sparse gradients in attention1e-4NoneYashudas et al. (2024)
EAWO-DNN (IoT)RMSpropα = 0.99, ε = 1e-8Better for non-stationary IoT data distributions1e-51.0 (max norm)Sornalakshmi et al. (2024)
HDNN (CNN-LSTM)Adam + SGD warmupAdam→SGD at epoch 50SGD generalizes better in late training1e-45.0 (LSTM)Reshan et al. (2023)
Embedded FS + DNNSGD + Momentummomentum = 0.9, nesterov = TrueStable convergence on small UCI dataset5e-4NoneZhang et al. (2021)
Physics-guided DLAdamβ₁ = 0.9, β₂ = 0.999Physics loss requires smooth gradient flow1e-31.0Zhang et al. (2023)
FL-LSTM (federated)SGD (local) + FedAvglr = 0.001 localSGD preferred in FL for communication efficiency1e-41.0Bojarczuk et al. (2024)
Transformer-CVDAdam (warmup)β₁ = 0.9, β₂ = 0.98Standard transformer training protocol0.1 (dropout-based)NoneBrowne et al. (2024)
NSGA-II ensembleAdam (per submodel)Default AdamEnsemble diversity from random initialization1e-4NoneGupta and Singh (2023)

Optimizer selection across CVD deep learning models.

4.6.4 Loss functions

Loss function analysis: The most complex imbalance-handling loss function reviewed is Focal Loss [used in FL-LSTM (Bojarczuk et al., 2024)], which down-weights easy negative training examples, and where the γ parameter (set to 2) serves to down-weight hard-to-classify minority-class (positive CVD) training cases. Transformer-CVD label smoothing is used to prevent overconfident predictions, using 0.05/0.95 instead of 0/1, to reduce calibration error and improve generalization (Browne et al., 2024). It is a physics-guided model that uses a composite loss combining two losses: a data-driven BCE and a physics residual loss that penalizes predictions outside the biologically plausible range. This is the most important distinguishing factor that enables the physics-guided approach to yield clinically plausible results (Zhang et al., 2023) (Table 31).

Table 31

ModelPrimary lossSecondary/auxiliary lossClass weightingImbalance strategyReferences
BiGRU-Attention (DEEP-CARDIO)Binary Cross-Entropy (BCE)Attention regularization lossInverse class frequencyNone (attention handles)Yashudas et al. (2024)
EAWO-DNN (IoT)Categorical cross-entropyEnergy consumption penaltyBalanced samplingESMO-based resamplingSornalakshmi et al. (2024)
HDNN (CNN-LSTM)Weighted BCEFeature consistency lossManual class weights (1:3 pos:neg)SMOTE pre-processingReshan et al. (2023)
Embedded FS + DNNBCEL1 sparsity on feature weightsNoneDataset balanced subsetZhang et al. (2021)
NSGA-II ensembleMulti-objective (accuracy + complexity)Pareto diversity lossNSGA-II guidedPareto front optimizationGupta and Singh (2023)
Physics-guided DLMSE (regression) + BCE (classification)Physics residual loss (PDE)N/APhysics prior constraintsZhang et al. (2023)
FL-LSTM + SHAP/LIMEFocal loss (γ = 2, α = 0.25)KL divergence (privacy)Focal dynamic weightingFocal loss handles imbalanceBojarczuk et al. (2024)
Transformer-CVDLabel smoothing cross-entropy (ε = 0.1)Auxiliary classification headsSmoothed labelsOversampling rare classesBrowne et al. (2024)

Loss function configurations across CVD models.

4.6.5 Software frameworks and hardware environments

Software environment observations: PyTorch and TensorFlow are equally represented - 51 vs. 49, respectively, among the reviewed studies, which is a reflection of the plurality of the field of software environments. The research-oriented studies (physics-guided, transformer, federated) are dominated by PyTorch because it supports a dynamic computational graph, enabling custom loss functions and manipulate gradients. Production-oriented IoT studies and clinical deployment studies are dominated by TensorFlow/Keras compatibility with edge hardware (Sornalakshmi et al., 2024; Yashudas et al., 2024). Hardware Gap: 31% of the studies reviewed failed to specify the hardware used to train the model, making it impossible to replicate the training time. The most performance-intensive model, Transformer-CVD on NVIDIA A100 (80GB) in 24 h, is a significant obstacle to the reproducibility of research groups that do not have access to a high-performance computing device (Table 32).

Table 32

ModelDeep learning frameworkLanguageKey librariesGPU hardwareRAMTraining timeReferences
BiGRU-Attention (DEEP-CARDIO)TensorFlow 2.x + KerasPython 3.8NumPy, Pandas, Scikit-learn, SHAPNVIDIA Tesla V100 (32GB)64GB~6 hYashudas et al. (2024)
EAWO-DNN (IoT)TensorFlow 2.xPython 3.7CloudSim 4.0, IoT-simNVIDIA GTX 1080 Ti (11GB)32GB~3 hSornalakshmi et al. (2024)
HDNN (CNN-LSTM)PyTorch 1.12Python 3.9torchvision, torchaudioNVIDIA A100 (40GB)128GB~12 hReshan et al. (2023)
Embedded FS + DNNKeras + TensorFlow 1.xPython 3.6Scikit-learn, SciPyNVIDIA GTX 1060 (6GB)16GB~1 hZhang et al. (2021)
Physics-guided DLPyTorch 1.10 + DeepXDEPython 3.8DeepXDE, FEniCSNVIDIA RTX 3090 (24GB)64GB~18 hZhang et al. (2023)
FL-LSTM + SHAP/LIMEPySyft + PyTorchPython 3.8PySyft 0.5, SHAP, LIME3 × NVIDIA GTX 1080 (distributed)32GB per node~8 h (federated)Bojarczuk et al. (2024)
Transformer-CVDPyTorch 1.13 + HuggingFacePython 3.10Transformers, TimmNVIDIA A100 (80GB)256GB~24 hBrowne et al. (2024)
NSGA-II ensembleTensorFlow + DEAPPython 3.7DEAP (EA library), KerasNVIDIA GTX 1080 (8GB)32GB~4 hGupta and Singh (2023)

Software and hardware environments across reviewed CVD studies.

4.6.6 Cross-validation and train/validation/test split protocols

Critical methodological gap—random seed reporting: Only 58% of the reviewed studies reported using fixed random seeds, which is a fundamental requirement of reproducibility. Stochastic components (weight initialization, dropout masks, data shuffling) lack a fixed seed, thereby contributing to non-deterministic behavior across runs and making it impossible to exactly replicate the results. This review suggests seed reporting (training seed, data-split seed, augmentation seed) as a minimum standard of reproducibility of CVD AI publications (Singh et al., 2024; Rahman et al., 2024). Gap of External validation: Only 12 of 69 studies reviewed (17%) conducted external validation using held-out datasets across various institutions. The remaining 83% report only internal cross-validation performance, which is systematically overestimated by 3–8% in AUC, according to meta-analytic estimates (Singh et al., 2024; Banerjee et al., 2025) (Table 33).

Table 33

ModelSplit strategyTrain %Val %Test %CV foldsExternal validationSeed fixedReferences
DEEP-CARDIO BiGRUStratified k-fold7015155-foldNoYes (seed = 42)Yashudas et al. (2024)
EAWO-DNNRandom split801010NoneNoNot reportedSornalakshmi et al. (2024)
HDNN CNN-LSTMStratified k-fold70151510-foldYes (holdout)Yes (seed = 0)Reshan et al. (2023)
Embedded FS + DNNHold-out8020NoneNoNot reportedZhang et al. (2021)
Physics-guided DLStratified k-fold7015155-foldNoYesZhang et al. (2023)
FL-LSTMPer-site local split701020None3-site holdoutYesBojarczuk et al. (2024)
Transformer-CVDStratified k-fold7510155-foldNoYes (seed = 123)Browne et al. (2024)

Data Split and Cross-Validation Protocols.

4.6.7 Methodological reporting completeness—gap analysis

Gap analysis narrative: The methodological reporting completeness audit indicates that the most important gaps in reproducibility in the reviewed CVD DL literature involve the external validation (17%) and the public code availability (12%). Even though studies report most training parameters, external validation is not established, so the reported benchmark performance figures are likely to be a clinically significant overestimation of actual clinical utility in the real world. This review is the first to recommend external validation, report the random seed, and release code as unnegotiable conditions for future CVD AI publications (Singh et al., 2024; Rahman et al., 2024; Ayoub et al., 2025) (Table 34).

Table 34

Methodological elementStudies reporting (n)Reporting rateReproducibility impactRecommendation
Optimizer specified52/6975%HighMandatory
Learning rate reported48/6970%HighMandatory
Batch size reported51/6974%MediumMandatory
Epochs / training rounds44/6964%HighMandatory
Loss function stated51/6974%Very highMandatory
Dropout rate reported39/6957%HighMandatory
GPU hardware specified47/6968%MediumRecommended
Software framework + version53/6977%HighMandatory
Random seed fixed and reported40/6958%Very highMandatory
External validation performed12/6917%Very highMandatory for clinical translation
Calibration metrics reported31/6945%Very highMandatory for clinical use
Code/model publicly available8/6912%CriticalStrongly recommended

Methodological reporting completeness across 69 reviewed studies.

5 Future research directions—accelerating clinical translation of multimodal CVD AI

Although benchmark accuracies exceed 99%, only 3 of 69 reviewed studies (4.3%) report prospective clinical validation, and none report regulatory approval or active clinical deployment. Table 35 summarizes 23 specific research opportunities in five critical directions: multimodal integration, real-time deployment, federated learning, longitudinal studies, and prospective clinical trials. Each has its current gap, proposed approach, expected clinical impact, and implementation timeline.

Table 35

DirectionResearch opportunityCurrent gapProposed approachClinical impactTimelineReferences
Multimodal integrationPan-modal fusion (6 + modalities simultaneously)No reviewed model fuses >4 modalitiesGraph neural networks with modality-specific encodersVery high3–5 yrsBrowne et al. (2024); Yang et al. (2024)
Temporal alignment of heterogeneous sampling ratesECG (360 Hz) vs. labs (daily) not synchronizedTemporal alignment networks with learnable resamplingVery high2–3 yrsYashudas et al. (2024); An et al. (2019)
Missing modality robustness via generative imputationEarly fusion collapses with any missing modalityVAE/diffusion model synthesis of missing streamsVery High2–3 yrsAn et al. (2019); Alasmari et al. (2025)
Multi-omics graph integration (proteomics + genomics)Omics is rarely used in DL CVD pipelinesGraph Convolutional Networks on biological pathwaysHigh4–6 yrsKrishnan et al. (2021); Wong and Tse (2021)
Real-time deploymentModel compression for wearable edge devicesTop models require a V100 GPU (32GB VRAM)Knowledge distillation + INT8 quantizationVery high1–3 yrsSornalakshmi et al. (2024); Yashudas et al. (2024)
Latency reduction for bedside emergency alertsTransformer inference: 200–500 ms per patientTensorRT optimization + structured pruningVery high1–2 yrsYashudas et al. (2024); Browne et al. (2024)
Battery-efficient continuous IoT monitoringSensors drain in <12 h under DL inferenceEAWO-DNN ESMO clustering on ARM Cortex-M chipsHigh2–3 yrsSornalakshmi et al. (2024)
5G cloud-edge hybrid for emergency inferenceHigh accuracy requires cloud; latency is too highSplit computing: edge preprocessing, cloud inferenceHigh3–4 yrsAlasmari et al. (2025); Otoum et al. (2024)
Federated learningDifferential privacy + XAI simultaneouslyFL-LSTM achieves privacy OR XAI—not both formallyDP-SGD with SHAP sensitivity analysis (ε < 1.0)Very high2–4 yrsBojarczuk et al. (2024)
Cross-continental FL networks (global equity)All FL studies use ≤3 sites, single countryFHIR-standardized international FL consortiaVery high4–7 yrsAlasmari et al. (2025); Otoum et al. (2024)
Personalized federated learning per patientFedAvg produces one global model onlyFedPer/MAML meta-learning for per-patient tuningHigh3–5 yrsOtoum et al. (2024); Bojarczuk et al. (2024)
Heterogeneous FL for non-IID multi-site dataMulti-site CVD data is highly non-IIDFedProx/SCAFFOLD for non-IID convergenceHigh2–3 yrsAlasmari et al. (2025); Otoum et al. (2024)
Longitudinal studiesContinuous wearable CVD trajectory cohortAll 69 reviewed studies are cross-sectionalProspective cohort: wearable + EHR, n ≥ 5,000, 5–10 yrsVery high5–10 yrsYashudas et al. (2024); Yang et al. (2024)
Treatment response monitoring (pre/post medication)No model tracks CVD risk change with treatmentPre/post multimodal monitoring with DL trajectory modelVery high2–5 yrsReshan et al. (2023); Zhang et al. (2023)
Genomic risk score evolution over timeStatic polygenic scores ignore gene–environment changesAnnual methylation + proteomics + imaging profilingHigh5–7 yrsKrishnan et al. (2021); Wong and Tse (2021)
Early life CVD risk seeding (age 20–35 yrs)Young adults are absent from all reviewed datasetsYoung-adult inception cohort, 20–30 year follow-upHigh10–20 yrsSingh et al. (2024); Yang et al. (2024)
Prospective clinical trialsPhase I—safety and feasibility (n = 100–500)0 of 69 reviewed studies conducted any clinical trialSingle-arm observational; AI prediction + physician decisionCritical1–2 yrsSingh et al. (2024); Ayoub et al. (2025)
Phase II—efficacy RCT (n = 500–2,000)No AI vs. standard care CVD trial existsRandomized: AI-assisted vs. standard care; 12–24 monthsCritical3–5 yrsSingh et al. (2024); Ayoub et al. (2025)
Phase III—outcomes RCT (n = 5,000–20,000)No MACE outcome trial for multimodal CVD AIDouble-blind RCT; primary endpoint: time-to-first MACECritical5–8 yrsSingh et al. (2024); Ayoub et al. (2025)
Equity-stratified trial (sex, ethnicity, SES strata)FN rate 4.8% higher in female patients (Section 4.5.5)Power trial to detect differential performance by subgroupVery High3–5 yrsSingh et al. (2024); Bojarczuk et al. (2024)
Regulatory translationFDA breakthrough device designation0 reviewed models have applied for FDA clearancePost Phase II evidence + pre-submission FDA meetingVery High4–6 yrsSingh et al. (2024); Ayoub et al. (2025)
TRIPOD-AI + CONSORT-AI complianceOnly 31% of reviewed studies are TRIPOD-AI compliantAdopt as a mandatory editorial and submission standardVery HighImmediateSingh et al. (2024); Rahman et al. (2024)
CE Mark under MDR 2017/745 (Europe)No reviewed multimodal CVD AI model CE markedClinical evaluation report + post-market surveillanceHigh4–7 yrsAyoub et al. (2025)

Consolidated future research directions for multimodal CVD AI.

5.1 Multi-modal integration

Temporal alignment of heterogeneous sampling rates—ECG at 360 Hz, wearables at 1 Hz, labs updated daily, and imaging annually cannot just be concatenated. The highest-priority architectural innovations for real-world clinical applications, where 30–60% of patients have not recorded their data, are temporal alignment networks and missing modality robustness via VAE/diffusion imputation (An et al., 2019; Alasmari et al., 2025).

5.2 Real-time deployment

The critical bottleneck is the accuracy-latency-power trilemma: to be highly accurate, the model must consume substantial GPU resources, which are incompatible with the constraints of wearables. Distillation of knowledge retains 95–97 thousand teacher accuracy at a cost 10 times lower (Sornalakshmi et al., 2024; Yashudas et al., 2024). The regulatory process takes five steps, one after another, where (1) IRB-approved prospective data collection, (2) TRIPOD-AI validation, (3) FDA pre-submission meeting, (4) De Novo/510 (k) application, (5) post-market surveillance; a 4–7 year gap between research and deployment.

5.3 Federated learning

Transatlantic FL networks between hospital systems in North America, Europe, Asia, and Africa would essentially fix the demographic bias issue cited in Section 4.5.5, which is currently barred by GDPR/HIPAA jurisdictional issues. The long-term vision of personalized federated learning via meta-learning (FedPer/MAML) is one in which a continually updated personal CVD risk model is maintained by each patient’s wearable (Otoum et al., 2024; Bojarczuk et al., 2024).

5.4 Longitudinal studies

It is the most important gap in the reviewed literature—all 69 papers are cross-sectional or have a short time window. The longitudinal CVD AI changes the clinical question from whether this patient is at high risk to when this patient is and is not at high risk. It necessitates ongoing wearable monitoring platforms, longitudinal EHR linkage, and time-aware architectures, such as Neural ODEs and temporal transformers (Yashudas et al., 2024; Browne et al., 2024; Yang et al., 2024). The primary source data is the ongoing repeat-imaging study by the UK Biobank.

5.5 Prospective clinical trials

None of the 69 reviewed studies was a prospective RCT, the largest obstacle to clinical adoption. The most urgently needed are three trial designs (1) alert fatigue trial that will test whether AI CVD alerts are more effective in improving physician response than dangerous over-alerts; (2) equity-stratified trial powered to detect superior performance across sex, ethnicity and socioeconomic strata; and (3) federated multi-site RCT that will simultaneously generate clinical evidence and advance the FL methodology (Singh et al., 2024; Ayoub et al., 2025) (Table 36).

Table 36

Research directionClinical impactFeasibilityTimelinePriority ★
Missing modality robustnessVery highHigh2–3 yrs★★★★★
Prospective Phase II clinical trialCriticalMedium3–5 yrs★★★★★
Federated learning + DP + XAIVery highMedium2–4 yrs★★★★
Longitudinal wearable cohortVery highMedium5–10 yrs★★★★
Edge/wearable model compressionHighHigh1–3 yrs★★★★
Sex and ethnicity equity trialVery highHigh3–5 yrs★★★★
TRIPOD-AI/CONSORT-AI complianceVery highHighImmediate★★★★
FDA regulatory submissionCriticalLow5–8 yrs★★★
Pan-modal genomics + imaging fusionHighLow4–6 yrs★★★
Cross-continental FL consortiumVery highLow5–8 yrs★★

Future research priority matrix—impact vs. feasibility.

Thangaraj et al. (2024) in the European Heart Journal, demonstrated that digital twin technology, a combination of in silico replication of patients with generative AI and multimodal data streams, enables personalized CVD risk simulation and evaluation of a virtual treatment scenario, a paradigm shift from the reactive cardiology paradigm.

6 Conclusion and future suggestions

Multimodal deep learning models, particularly hybrid ensemble architectures and BiGRU-Attention frameworks, have set a new benchmark for CVD risk stratification in both IoT-enabled and clinical settings, consistently achieving over 98% accuracy, precision, and recall. Innovations such as ESMO energy-aware clustering, NSGA-II feature selection, and CNN-GRU-LSTM hybrid architectures effectively address challenges related to data heterogeneity, computational efficiency, and personalized care, thereby enabling continuous monitoring, earlier intervention, scalable precision medicine, and more patient-centered healthcare delivery. Looking ahead, the field must prioritize external validation in diverse, prospective multicenter cohorts to ensure robust generalizability, especially for underrepresented populations; implement federated learning frameworks with differential privacy to enable secure multi-institutional deployment; embed physics-guided modeling and attention-based interpretability directly into model architectures to enhance transparency and clinician trust; and adopt TRIPOD-AI and CONSORT-AI reporting standards as essential requirements for publication and regulatory approval. Additional priorities include developing continuous learning systems that adapt to incoming patient data without catastrophic forgetting, integrating multi-omics data using graph neural networks and multimodal transformers to uncover hidden causal pathways, and advancing digital twin frameworks that merge physiological knowledge with data-driven intelligence for auditable, patient-specific forecasting and virtual clinical trials. Collectively, evidence from 69 studies published between 2012 and 2025, including DEEP-CARDIO BiGRU-Attention (99.9% accuracy), EAWO-DNN for IoT applications (98.9% accuracy, 0.99 AUC), and FL-LSTM with explainable AI (99% AUC across three ECG datasets), demonstrates the rapid evolution of this field. Notably, 59% of these studies were published in 2024–2025, underscoring the accelerating trajectory toward precision cardiovascular medicine.

Statements

Author contributions

GN: Writing – original draft, Writing – review & editing, Methodology, Conceptualization, Data curation. MG: Writing – original draft, Writing – review & editing, Conceptualization, Data curation, Resources, Supervision, Validation, Visualization.

Funding

The author(s) declared that financial support was received for this work and/or its publication. Open access funding was provided by Vellore Institute of Technology, Vellore, India.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that Generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/frai.2026.1840804/full#supplementary-material

References

  • 1

    AddissoukyT. A.HarbH. A.MohamedS. H.El-SayedA. S.SamyM. (2024). Recent developments in the diagnosis, treatment, and management of cardiovascular diseases through artificial intelligence and other innovative approaches. PRO5, 2940. doi: 10.46439/biomedres.5.041,

  • 2

    AhmadM.AlfayadM.AftabS.KhanM. A.FatimaA.ShoaibB.et al. (2021). Data and machine learning fusion architecture for cardiovascular disease prediction. Comput. Mater. Contin.69:13. doi: 10.32604/cmc.2021.019013

  • 3

    AhujaS. K.ShrimankarD. D.DurgeA. R.Design of an Iterative Method for deep multimodal feature fusion in heart disease diagnostics utilizing explainable AI, in Proceeding 1st International Conference Explainable AI for Neural and Symbolic Methods (EXPLAINS 2024), (2024), pp. 8795.

  • 4

    AlasmariS.AlGhamdiR.TejaniG. G.Kumar SharmaS.MousaviradS. J. (2025). Federated learning-based multimodal approach for early cardiac disease detection. Front. Cardiovasc. Med.16:1563185. doi: 10.3389/fphys.2025.1563185,

  • 5

    AlmulihiA. H.AlmulihiA.SalehH.HussienA. M.MostafaS.El-SappaghS.et al. (2022). Ensemble learning based on hybrid deep learning model for heart disease early prediction. Diagnostics12:215. doi: 10.3390/diagnostics12123215,

  • 6

    AlzahraniA.RazaM. A.AsgharM. Z. (2025). Demystifying diagnosis: An efficient deep learning technique with explainable AI to improve breast Cancer detection. PeerJ Comput. Sci.11:e2806. doi: 10.7717/peerj-cs.2806,

  • 7

    AmalS.SawyerD.KönikA. (2024). Artificial intelligence and multimodal medical imaging data fusion for improving cardiovascular disease care. Front. Radiol.4:1412404. doi: 10.3389/fradi.2024.1412404,

  • 8

    AnY.HuangN.ChenX.WuF. X.WangJ. (2019). High-risk prediction of cardiovascular diseases via attention-based deep neural networks. IEEE/ACM Trans. Comput. Biol. Bioinform.18, 10931105. doi: 10.1109/TCBB.2019.2935059,

  • 9

    AyoubC.CamusY.RoussetJ.DercleD.RaymondD.BosqJ.et al. (2025). Multimodal fusion artificial intelligence model to predict myocarditis and adverse cardiovascular events in cancer patients starting ICI therapy. JACC Advances4:101435. doi: 10.1016/j.jacadv.2024.101435,

  • 10

    BanerjeeT.ChhabraP.KumarM.KumarA.AbhishekK.ShahM. A. (2025). Pyramidal attention-based T network for brain tumor classification: a comprehensive analysis of transfer learning approaches for clinically reliable and reliable AI hybrid approaches. Sci. Rep.15:28669. doi: 10.1038/s41598-025-11574-x,

  • 11

    BarbieriS.MehtaS.WuB.BharatC.PoppeK.JormL.et al. (2020). Predicting cardiovascular risk from national administrative databases using a combined survival analysis and deep learning approach. Int. J. Epidemiol.51, 931944. doi: 10.1093/ije/dyab258,

  • 12

    BhagawatiM.PaulS.AgarwalS.ProtogeronA.SfikakisP.KitasG.et al. (2023). Cardiovascular disease/stroke risk stratification in deep learning framework: a review. Cardiovasc. Diagn. Therapy12, 557598. doi: 10.21037/cdt-22-438,

  • 13

    BojarczukV.SpinosaL. N.de CarvalhoA. C.DietrichJ. L. G. (2024). Hybrid explainable deep learning for heart disease prognosis. Appl. Soft Comput.141.

  • 14

    BrowneS.et al. (2024). Transformer-based attention models for cardiovascular risk. Sci. Rep.14.

  • 15

    CocianuC.UscatuC.KofidisK.MuraruS.VăduvaA. (2023). Classical, evolutionary, and deep learning approaches of automated heart disease prediction: a case study. Electronics12:663. doi: 10.3390/electronics12071663

  • 16

    DehghaniS.et al. (2025). Explainable AI in cardiovascular health: multimodal analysis and prediction. IEEE Rev. Biomed. Eng.18, 110131.

  • 17

    DharmarathneG.BogahawaththaM.RathnayakeU.MeddageD. P. P. (2024). Integrating explainable machine learning and user-centric design for cardiovascular disease prediction. Patterns23:200428. doi: 10.1016/j.iswa.2024.200428,

  • 18

    DuY.HuangK.YueM.ChenC.DengX.WangN. (2026). CASNet: curvature-aware cardiac MRI segmentation with multi-scale and attention-driven encoding for enhanced risk-oriented structural analysis. Front. Med.12:1688872. doi: 10.3389/fmed.2025.1688872,

  • 19

    El-SofanyH.El-SayedA.TawfikA. H. (2024). A proposed technique for predicting heart disease using feature selection. Sci. Rep.14:23277. doi: 10.1038/s41598-024-74656-2,

  • 20

    Evangelin SoniaS. V.JuliusA.MadusudhanaR. (2023). A multi-modal integrated deep neural network for the prediction of cardiovascular disease. Autom. Remote. Control.

  • 21

    FathimaA. J.FaslaM. M. N. (2024). A comprehensive review on heart disease prognostication using different artificial intelligence algorithms. Comput. Methods Biomech. Biomed. Engin.27, 13571374. doi: 10.1080/10255842.2024.2319706,

  • 22

    GaoJ.LiP.ChenZ.ZhangJ. (2020). A survey on deep learning for multimodal data fusion. Neural Comput.32, 829864. doi: 10.1162/neco_a_01273,

  • 23

    GuptaA.SinghA. (2023). EDL-NSGA-II: ensemble deep learning framework with NSGA-II feature selection for heart disease prediction. Expert. Syst.40:254. doi: 10.1111/exsy.13254

  • 24

    HaqI.LiangH.ZengK.WangT.UddinI.LinJ.et al. (2026). Deep learning advancements for cardiovascular diseases (CVDs) diagnosis: imaging modalities, challenges, and future perspectives. Biomed. Signal Process. Control119:109899. doi: 10.1016/j.bspc.2026.109899

  • 25

    HaqI.et al. (2020). Artificial intelligence in personalized cardiovascular medicine and cardiovascular imaging. Cardiovasc. Diagn. Ther.11, 911923. doi: 10.21037/cdt.2020.03.09,

  • 26

    HathawayQ. A.YanamalaN.BudoffM. J.SenguptaP. P.ZebI. (2021). Deep neural survival networks for CVD risk prediction: MESA. Comput. Biol. Med.139:104983. doi: 10.1016/j.compbiomed.2021.104983,

  • 27

    HoltA.BatinicaB.LiangJ.KerrA.CrengleS.HudsonB.et al. (2023). Development and validation of cardiovascular risk prediction equations in 76,000 people with known cardiovascular disease. Eur. J. Prev. Cardiol.31:314. doi: 10.1093/eurjpc/zwad314

  • 28

    HoniD. G.OkekeochaS.EdehT. T.JacksonA.EzenwakaC. U. (2024). A one-dimensional convolutional neural network-based method for cardiovascular disease prediction. Inform. Med. Unlocked49:101535. doi: 10.1016/j.imu.2024.101535

  • 29

    JaltotageB.ChandraratneA.LuciniS.MaharZ. U.RiederM. J.NorgroveL.et al. (2024). Use of artificial intelligence including multimodal systems in cardiovascular disease management. Canad. J. Cardiol. Open40, 18041812. doi: 10.1016/j.cjca.2024.07.014,

  • 30

    KheraR.AlexanderK. P.CutlerM. J.DandekarS. P.BrosiusE. N.FeldmanH. I. (2024). Transforming cardiovascular care with artificial intelligence: from discovery to practice. J. Am. Coll. Cardiol.84, 97114. doi: 10.1016/j.jacc.2024.05.003,

  • 31

    KrishnaR.GuptaM.SharmaD.GoyalS. (2023). Ensemble deep learning for multimodal heart disease detection. IEEE Access.

  • 32

    KrishnanS.MagalingamP.IbrahimR. (2021). Hybrid deep learning model using recurrent neural network and gated recurrent unit for heart disease prediction. Int. J. Electr. Comput. Eng.11, 54675476. doi: 10.11591/ijece.v11i6.pp5467-5476

  • 33

    KrittanawongC.JohnsonK. W.RosensonR. S.WangZ.AydarM.BaberU.et al. (2019). Deep learning for cardiovascular medicine: a practical primer. Eur. Heart J.40, 20582073. doi: 10.1093/eurheartj/ehz056,

  • 34

    LiY.LiH.MeiF.XiongQ. (2024). A review of deep learning-based information fusion for medical multimodal classification tasks. Comput. Biol. Med.177:108635. doi: 10.1016/j.compbiomed.2024.108635,

  • 35

    LiY.LiK.WangB.LiuX. (2025). Heart failure prognosis risk assessment model based on multimodal data fusion. Egypt. Inform. J.26.

  • 36

    LiX.LiuX.DengX.FanY. (2022). Interplay between artificial intelligence and biomechanics modeling in the cardiovascular disease prediction. Biomedicine10:157. doi: 10.3390/biomedicines10092157,

  • 37

    MederB.VölzkeC.MaasL. R.MüllerS. E.ThormannT. T.FelzenF. T.et al. (2025). Artificial intelligence to improve cardiovascular population health: state-of-the-art and future directions. Front. Cardiovasc. Med.46, 19071916. doi: 10.1093/eurheartj/ehaf125,

  • 38

    MulaniA.et al. (2025). ML-powered internet of medical things structure for heart disease prediction. J. Pharmacol. Pharmacother.16, 3845. doi: 10.1177/0976500X241281490

  • 39

    NarayanY.SinghD. P.BanerjeeT.KourP.RaneK.CA. D. D.et al. (2026). A comparative evaluation of deep learning architectures for prostate cancer segmentation: introducing TrionixNet with N-Core multi-attention mechanism. Arch. Comput. Methods Eng.33, 37073746. doi: 10.1007/s11831-025-10411-8

  • 40

    OgunpolaA.BakareS.OgunpolaO. (2024). Machine learning-based predictive models for detection of heart disease. Diagnostics14:144. doi: 10.3390/diagnostics14020144,

  • 41

    OhS.ShimJ. Y. (2024). Development and validation of a deep learning–based cardiovascular disease risk prediction model for long-term breast cancer survivors. J. Clin. Oncol.42:12023. doi: 10.1200/JCO.2024.42.16_suppl.12023

  • 42

    OrdikhaniM.AbadehM. S.PruggerC.HassannejadR.MohammadifardN.SarrafzadeganN. (2022). An evolutionary machine learning algorithm for cardiovascular disease risk prediction. PLoS One17:723. doi: 10.1371/journal.pone.0271723,

  • 43

    OtoumS.et al. (2024). Federated deep learning for smart healthcare IoT: a privacy-preserving CVD screening model. IEEE Internet Things J.

  • 44

    RahmanS.et al. (2024). Multimodal data fusion for precision cardiology: current trends and future prospects. BMC Med. Inform. Decis. Mak.

  • 45

    ReshanM. S. A.AminS.ZebM. A.SulaimanA.AlshahraniH.ShaikhA. (2023). A robust heart disease prediction system using hybrid deep neural networks. IEEE Access11, 121574121591. doi: 10.1109/access.2023.3328909

  • 46

    SadrH.KhalilzadehM. R.RezaeiA.AshrafiN.ZakeriP. (2024). Cardiovascular disease diagnosis: a holistic approach using deep learning and machine learning. Front. Cardiovasc. Med.29:44. doi: 10.1186/s40001-024-02044-7,

  • 47

    SahaP.et al. (2023). A comprehensive survey on explainable multimodal data fusion for CVD risk stratification. Artif. Intell. Rev.

  • 48

    SekarB.DongM.ShiJ.HuX. (2012). Fused hierarchical neural networks for cardiovascular disease diagnosis. IEEE Sensors J.12, 644650. doi: 10.1109/JSEN.2011.2129506

  • 49

    SelvarathiC.VaradhaganapathyS. (2023). Deep learning based cardiovascular disease risk factor prediction among type 2 diabetes mellitus patients. Inf. Technol. Control52, 215227. doi: 10.5755/j01.itc.52.1.32008

  • 50

    ShaikT.ReddyB. K.KumarA.KumarP. (2024). A survey of multimodal information fusion for smart healthcare. Inf. Fusion102:102040. doi: 10.1016/j.inffus.2023.102040,

  • 51

    SinghD. P.BanerjeeT.KourP.MalikR.NaiduG. R.CA. D. D.et al. (2026a). A comprehensive study of enhanced computational approaches for breast cancer classification: comparative analysis with existing state of the art methods. Arch. Comput. Methods Eng.33, 36353663. doi: 10.1007/s11831-025-10414-5

  • 52

    SinghD. P.BanerjeeT.MahajanS.Ramesh ChandraK.KumarR.PhaniS.et al. (2026b). A comprehensive study on deep learning models for the detection of diabetic retinopathy using pathological images. Arch. Comput. Methods Eng.33, 503532. doi: 10.1007/s11831-025-10315-7

  • 53

    SinghM.GuptaA.VermaD.TiwariN.LodhaJ. (2024). Artificial intelligence for cardiovascular disease risk prediction: review and future directions. J. Cardiovasc. Comput. Tomogr.

  • 54

    SornalakshmiM.DevakanthJ. J. M. A.RajalakshmiR.VelmurugadassP. (2024). Energy-aware heart disease prediction using ESMO and optimal deep learning model for healthcare monitoring in IoT. J. Biomol. Struct. Dynamics43, 35423556. doi: 10.1080/07391102.2023.2298736,

  • 55

    StahlschmidtS.et al. (2022). Multimodal deep learning for biomedical data fusion: a review. Brief. Bioinform.23:569. doi: 10.1093/bib/bbab569,

  • 56

    SumalathaU.PrakashaK.PrabhuS.NayakV. C. (2024). Deep learning applications in ECG analysis and disease detection: an investigation study of recent advances. IEEE Access12, 126258126284. doi: 10.1109/access.2024.3447096

  • 57

    ThangarajP. M.BensonS. H.OikonomouE. K.AsselbergsF. W.KheraR. (2024). Cardiovascular care with digital twin technology in the era of generative artificial intelligence. Eur. Heart J.45, 48084821. doi: 10.1093/eurheartj/ehae619,

  • 58

    VenkateshC.PrakashS.VenkatP.NarayananR. (2024). An automatic diagnostic model for detection and prognosis of cardiovascular disease using multimodal health record data. Heliyon10:e25574. doi: 10.1016/j.heliyon.2024.e25574,

  • 59

    VincentP. R.et al. (2022). IoT-cloud-based smart healthcare monitoring system for heart disease prediction via deep learning. Electronics.

  • 60

    WangJ.DingH.BidgoliF. A.ZhouB.IribarrenC.MolloiS.et al. (2017). Detecting cardiovascular disease from mammograms with deep learning. IEEE Trans. Med. Imaging36, 11721181. doi: 10.1109/TMI.2017.2655486,

  • 61

    WangY.WuL.ChenZ.NiX.SuM. (2024). Cardiovascular disease prediction model based on patient behavior patterns in deep learning. Front. Psych.15:1418969. doi: 10.3389/fpsyt.2024.1418969,

  • 62

    WangG.ZhangY.LiS.ZhangJ.JiangD.LiX.et al. (2021). A machine learningbased prediction model for cardiovascular risk in women with preeclampsia. Front. Cardiovasc. Med.8:736491. doi: 10.3389/fcvm.2021.736491,

  • 63

    WaniN. A.KumarR.BediJ. (2024). DeepXplainer: An interpretable deep learning based approach for lung Cancer detection using explainable artificial intelligence. Comput. Methods Prog. Biomed.243:107879. doi: 10.1016/j.cmpb.2023.107879,

  • 64

    WehbeR. M.KatsaggelosA. K.HammondK. J.HongH.AhmadF. S.OuyangD.et al. (2023). Deep learning for cardiovascular imaging: a review. JAMA Cardiol.8:3142. doi: 10.1001/jamacardio.2023.3142,

  • 65

    WongY.TseH. (2021). Circulating biomarkers for cardiovascular disease risk prediction in patients with cardiovascular disease. Front. Cardiovasc. Med.8:713191. doi: 10.3389/fcvm.2021.713191,

  • 66

    YangH.HuangQ.DuanX.HuangL.WangM. (2024). AI-powered CVD risk modeling from longitudinal multimodal imaging. Front. Radiol.

  • 67

    YashudasA.GuptaD.PrashantG. C.DuaA.AlQahtaniD.ReddyA. S. K. (2024). DEEP-CARDIO: recommendation system for cardiovascular disease prediction using IoT network. IEEE Sensors J.24, 1453914547. doi: 10.1109/jsen.2024.3373429

  • 68

    YunH.NohN. I.LeeE. Y. (2022). Genetic risk scores used in cardiovascular disease prediction models: a systematic review. Rev. Cardiovasc. Med.23:8. doi: 10.31083/j.rcm2301008,

  • 69

    ZhangD.ChenY.YeS.CaiW.JiangJ.XuY.et al. (2021). Heart disease prediction based on the embedded feature selection method and deep neural network. J. Healthcare Eng.2021:22. doi: 10.1155/2021/6260022,

  • 70

    ZhangD.LiuX.XiaJ.GaoZ.ZhangH.de AlbuquerqueV. H. C. (2023). A physics-guided deep learning approach for functional assessment of cardiovascular disease in IoT-based smart health. IEEE Internet Things J.10, 1850518516. doi: 10.1109/jiot.2023.3240536

  • 71

    ZhuJ. Y.ZhangL.XiaoS.XuZ.TanQ., Cardiovascular disease detection based on multi-modal data and multi-branch residual networks, in Proceeding ACM International Conference, (2023).

  • 72

    ZhuJ.et al. (2025). Dual-scale deep residual network for multimodal cardiovascular disease detection. Comput. Biol. Med.171.

Appendix: transparency, reproducibility, and open science statement

In accordance with open science principles and reviewer guidance, this systematic review provides the following transparency and reproducibility resources:

  • PRISMA protocol and search reproducibility: The complete literature search strategy—including Boolean query strings, database sources (PubMed, Scopus, IEEE Xplore, Web of Science, Google Scholar), inclusion/exclusion criteria, and PRISMA 2020 flow diagram—is fully documented in Section 1.1, enabling independent replication of the search and screening process.

  • Data extraction sheet: A structured data extraction template covering model architecture, datasets, performance metrics, fusion strategy, and XAI method used for all 69 included studies is provided as Supplementary File accompanying this manuscript. The template organises 64 extraction sub-fields across seven thematic groups (Study Identification, Data Sources & Modalities, Methodology & Architecture, Training & Validation, Performance & Explainability, Clinical Translation & Bias, and Quality & Risk of Bias), enabling independent replication and extension of the extraction process.

  • Reproducibility of reviewed models: This review synthesizes findings from 69 peer-reviewed studies. For original model code and pretrained weights, readers are directed to the primary source repositories cited in each study. Key models with publicly available code include:

  • Living review commitment: This review will be updated as significant new studies emerge. Version-controlled updates will be tracked via the GitHub repository with tagged releases corresponding to manuscript versions.

Summary

Keywords

BiGRU-Attention, cardiovascular disease, deep learning, explainability, federated learning, IoT healthcare, multimodal data fusion, risk prediction

Citation

Ganeshan N and Magesh G (2026) Deep learning for cardiovascular disease: a comprehensive review of detection and risk forecasting. Front. Artif. Intell. 9:1840804. doi: 10.3389/frai.2026.1840804

Received

27 March 2026

Revised

05 May 2026

Accepted

22 May 2026

Published

26 June 2026

Volume

9 - 2026

Edited by

Caetano Mazzoni Ranieri, São Paulo State University, Brazil

Reviewed by

Niyaz Ahmad Wani, Manipal University Jaipur, India

Tathagat Banerjee, Indian Institute of Technology Patna, India

Updates

Copyright

*Correspondence: G. Magesh,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics