ORIGINAL RESEARCH article

Front. Neuroinform., 28 March 2022

Volume 16 - 2022 | https://doi.org/10.3389/fninf.2022.856295

Research on Voxel-Based Features Detection and Analysis of Alzheimer’s Disease Using Random Survey Support Vector Machine

  • 1. School of Computer Information and Engineering, Changzhou Institute of Technology, Changzhou, China

  • 2. School of Computer Science and Engineering, Changshu Institute of Technology, Changshu, China

  • 3. School of Computer Science and Artificial Intelligence, Changzhou University, Changzhou, China

Abstract

Alzheimer’s disease (AD) is a degenerative disease of the central nervous system characterized by memory and cognitive dysfunction, as well as abnormal changes in behavior and personality. The research focused on how machine learning classified AD became a recent hotspot. In this study, we proposed a novel voxel-based feature detection framework for AD. Specifically, using 649 voxel-based morphometry (VBM) methods obtained from MRI in Alzheimer’s Disease Neuroimaging Initiative (ADNI), we proposed a feature detection method according to the Random Survey Support Vector Machines (RS-SVM) and combined the research process based on image-, gene-, and pathway-level analysis for AD prediction. Particularly, we constructed 136, 141, and 113 novel voxel-based features for EMCI (early mild cognitive impairment)-HC (healthy control), LMCI (late mild cognitive impairment)-HC, and AD-HC groups, respectively. We applied linear regression model, least absolute shrinkage and selection operator (Lasso), partial least squares (PLS), SVM, and RS-SVM five methods to test and compare the accuracy of these features in these three groups. The prediction accuracy of the AD-HC group using the RS-SVM method was higher than 90%. In addition, we performed functional analysis of the features to explain the biological significance. The experimental results using five machine learning indicate that the identified features are effective for AD and HC classification, the RS-SVM framework has the best classification accuracy, and our strategy can identify important brain regions for AD.

Introduction

Due to the development of medical technology, the world population has grown steadily, and the elderly population has increased rapidly. It is expected that this trend will continue to accelerate in the next few decades, and the occurrence of senile diseases and the social cost of aging are expected to increase. Alzheimer’s disease (AD) is a brain disease. It is also a progressive disease, meaning that it will get worse over time. It is believed that AD begins 20 years or more before the onset of symptoms (). The preclinical stage of AD is crucial for identifying early pathophysiological events and developing interventions for disease improvement. Given that changes in synaptic function occur early in the neurodegenerative process, functional MRI (fMRI) is particularly promising for detecting early changes in brain function (). MRI has aroused great interest in AD-related research due to its complete noninvasiveness, high availability, high spatial resolution, and good contrast between different soft tissues ().

Mild cognitive impairment (MCI), known as the early stage of AD, was a disease state of cognitive decline between normal elderly and dementia patients. MCI was divided into early mild cognitive impairment (EMCI) and late mild cognitive impairment (LMCI). Studies had pointed out that if MCI patients were not diagnosed early, the probability of developing AD could be as high as 80% after 6 years, and about two-thirds of AD patients were converted through MCI (; ; ). Using linear mixed models, analyzed 2,261 individuals with MCI and non-MCI and found that the neurodegeneration was associated with letter fluency and semantic fluency. introduced the linear regression classification to classify samples and obtained an accuracy of 97.51% (). ) applied the random forest to identify features associated with AD. Another study showed that the Flash Visual Evoked Potential-P2 latency had AD-specific pathological information (). calculated the degree of atrophy of hippocampus and cortical areas and found that the specific cortical thinning and the reduction of hippocampal volume were accelerated in early AD. As the classic analysis methods, the machine learning algorithms brought new research sight to AD-specific biomarkers (; ; ; ). applied a group lasso support vector machine to obtain the AD-specific biomarkers. developed two XGBoost classification models to classify AD and healthy control (HC). Studies have proved that AD was closely related to brain atrophy and that brain atrophy was mainly reflected in the reduction of cortical surface area, thickness, and gray matter volume and, therefore, gray matter volume, cortical surface area, and average thickness contributed to the pathology of AD patients (; ; ; ).

Despite many efforts, it is still challenging to determine effective AD-specific biomarkers for early diagnosis and prediction of disease progression and requires more research (; ). In our study, we proposed a novel analysis framework based on the Random Survey Support Vector Machines (RS-SVM) for the early detection of AD conversion in MCI patients by using advanced machine learning algorithms and combining voxel-based data with standard neuropsychological test results. First, to obtain the voxel sets, we extracted the differences between AD and HC. Then, we applied the RS-SVM to identify important features that classified EMCI, LMCI, AD, and HC well. Subsequently, we applied several classical methods to construct the analysis frameworks and evaluate the accuracy of these features to classify with EMCI-HC, LMCI-HC, and AD-HC. The experiment results demonstrate that the identified features were effective in classifying AD, the RS-SVM framework performed well, and the identified regions and genes will further our understanding of AD.

Materials and Methods

Figure 1 illustrates the framework of a voxel-based three-level analysis for AD. The framework encompasses data processing (A), features extraction (B), RS-SVM construction (C), and the gene-level analysis using effective chi-square statistic (ECS) method () and pathway-level analysis using the resulting genes () (D). The novation of this framework is to make the full use of voxel-based data.

FIGURE 1

Imaging Data

In this study, we downloaded and analyzed 1,426 participants with genotyping data and MRI scans from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database (adni.loni.usc.edu). These data include 353 HCs, 273 EMCI, 504 LMCI, and 296 with AD. The characteristics of these participants, including average age and years of education, are shown in Table 1.

TABLE 1

SubjectsHCEMCILMCIADp
Number353273504296
Gender (M/F)187/166153/120309/195166/130<0.001
Age (mean ± sd)72.2 ± 7.671.3 ± 7.174.0 ± 7.675.1 ± 5.5<0.001
Edu (mean ± sd)16.1 ± 2.716.1 ± 2.616.0 ± 2.916.3 ± 2.6<0.001

Participant characteristics.

HC, healthy control; EMCI, early mild cognitive impairment; LMCI, late mild cognitive impairment; AD, Alzheimer’s disease; Edu, education.

Random Survey Support Vector Machines-Based Machine Learning Method

Data Processing

MRI scans, using voxel-based morphometry (VBM), were aligned and normalized to a T1-weighted template image and the Montreal Neurological Institute (MNI) space, respectively. The gray matter density (GMD) maps were segmented, extracted, and smoothed with an 8-mm FWHM (full width at the half maximum) kernel. The Automatic Anatomical Labeling (AAL) atlas was employed to define the regions of interest (ROIs) and their coordinates (whole brain) (). We then down-sampled the resulting maps to a dimension of 61 × 73 × 61 to reduce the data size for subsequent analysis in EMCI-HC, LMCI-HC, and AD-HC groups.

To extract the differences within the three groups, i.e., A, B, and C, we first performed the weighted process of the two sets of images separately and saved them as matrices M and N (i.e., M for AD group and N for HC group). Then, let represents the vector of two voxels (), and we obtained a vector (k = 271, 633). Since the voxels () in the two groups were meaningless for our research, we deleted these voxels and obtained 64,411 sets of different voxels. We used V′ to denote the voxel set.

Feature Extraction

The data sets of features were still too many for our final binary classification in EMCI-HC, LMCI-HC, and AD-HC groups. Therefore, we estimated the number of features by calculating the similarity between two rows and in V′. The similarity of the voxels is given in the following equation:

where and (i,j = 1,2,…,64411) are the values of AD group. and (i,j = 1,2,…,64411) are the values of HC groups. ρi is the similarity between and .

For the convenience of calculation, we divided the V′ into ten groups and obtained 55 sets of similarity matrices. On this basis, we defined the number of minimal ρi as Cmin and the number of maximal ρi as Cmax. Due to the value of Cmin and Cmax (Cmin = 132, Cmax = 21), we defined that the number of features should be in [Cmax,Cmin]. Then, we extracted 64,411 features of all subjects from the original MR images to form a 649 × 64,411 matrix as the initial data set.

Random Survey Support Vector Machines Construction

To extract the important feature, we proposed a single-kernel SVM model based on random survey. The goal of random survey was the establishment of a random experimental data set. Since the initial data set X was a two-dimensional matrix of 649 × 64,411, we selected l column from the X randomly and constructed a single randomized experimental data set X′ (l ∈[Cmax,Cmin]). At the same time, the set of columns corresponding to each column l in each extraction was R = {r1,r2,…,rl}, which denoted the index of brain loci coordinates. The indices are extracted as follows:

After random extraction, we defined the training set:validation set:test set as 6:2:2. The training set was used as an input for training first. The validation set was applied to obtain the optimal hyperparameters and replaced the initial parameters. The remaining 20% was introduced as the test set to calculate the accuracy of the tuned model and to evaluate whether the obtained feature set R = {r1,r2,…,rl} can be used as the final feature set.

In the classification process of SVM, the input data and the learning objective y = {y1,y2,…,yN,…,yM} were given, where N was the number of EMCI, LMCI, and AD samples, respectively, and M was the number of HC. The learning objectives were binary variables y = {−1,1}, where -1 represents EMCI, LMCI, and AD, respectively, and 1 represents HC in the three groups. The feature set of the input data was regarded as the hyperplane D in decision boundaries to separate the learning targets by positive and negative classes, making the distance εi between any sample and plane ≥1. The hyperplane and the plane distance are defined as follows:

where w denotes the normal vector of the hyperplane and b denotes the intercept of the hyperplane. The decision boundary satisfying this condition actually constructed two parallel hyperplanes D1,D2 as interval boundaries to classify the samples (Eq. 4).

Based on Eq. 4, it could be derived that all samples above the upper interval boundary were positive and those below the lower interval boundary were negative. The distance between the two interval boundaries was defined as the margin. Since our experimental data X′ was selected randomly, there was hyperboloid in the feature set to separate positive and negative classes. Using nonlinear functions, the nonlinear separable problems from the original feature set were mapped to a higher dimensional Hilbert space H. The hyperplane, using as the decision boundary, is defined as follows:

where φ:X′↦H denotes the mapping function. Since the mapping function was complex, it was difficult to calculate the inner product. Therefore, the inner product of the mapping function was defined as kernel functions to avoid the explicit operation.

Parameter Determination

We used the original linear kernel function of the support vector machine first, and the penalty factor C and the kernel parameter gamma were set as default values (C = 1 and gamma = 0.5). Then, we applied the training data set and labels to train the model. Subsequently, the hyperparameters were optimized by grid search. The SVM could be transformed into an equivalent quadratic convex optimization problem to solve using the following equation:

Evaluation Metrics

In this article, the samples were positive and negative, and the results classified had the following cases:

True positive (TP): the positive sample was predicted as a positive sample.

True negative (TN): the negative sample was predicted as a negative sample.

False positive (FP): the negative sample was predicted as a positive sample.

False negative (FN): the positive sample was predicted as a negative sample.

Let P denotes the positive sample and N denotes the negative sample. We then obtained the following equation:

The evaluation metrics used in our research are as follows:

● Accuracy. Accuracy was the number of correctly classified samples divided by the total number of samples (Eq. 8).

● Precision. Precision was the proportion of the samples that were actually positive (or negative) divided by samples classified as positive (or negative) (Eq. 9).

● Recall. Recall was the measure of coverage (Eq. 10).

● Comprehensive evaluation indicators (F-Measure). Accuracy and sensitivity sometimes needed to be considered together as given in the following equation:

Whenα = 1, Eq. 11 is transformed into the following equation as follows:

Model Comparison

We used the test set to evaluate the classification ability of 5 machine learning methods, including linear regression model, least absolute shrinkage and selection operator (Lasso) model, partial least squares (PLS) model, SVM model, and RS-SVM model. First, the initial default parameters were applied to each model to train and calculate the evaluation metrics. Then, the grid search algorithm was used to optimize the hyperparameters of the five models. Finally, the hyperparameters were introduced in each model to recalculate the evaluation metrics. The results were used to evaluate the pros and cons of the five models.

Since the RS-SVM model in this article was optimized based on the traditional SVM model, the other three evaluation models were described in detail in this section.

Linear regression model was a statistical analysis method that used regression analysis in mathematical statistics to determine the quantitative relationship between the interdependence of two or more variables.

Given a data set D = {(x1,y1),(x2,y2),…,(xi,yi)}, we learned that a linear model from this data set will reflect the correspondence between xi and yi as accurately as possible. The linear regression model, which was a function of linear combination of attributes x, could be expressed as follows:

where W = {w1,w2,…,wi} is column vector, indicating the weight of the corresponding attribute in the prediction result. Eq. 13 was represented as the following equation:

Then it was to find a model such that ∀i ∈ [1,m] has f(xi) as close to yi. Therefore, the sum of the squares of the difference between the predicted value and the real value of each sample is minimized and thus gives the following equation:

where (w*,b*) is the optimal parameter, and the minimum value of (w,b) is taken for the above equation.

The Lasso model was a compression estimation method with the idea of reducing the variable set (decreasing order). By constructing a penalty function, it could compress the coefficients of variables and made some regression coefficients become 0, so as to achieve the purpose of variable selection.

Given n data samples {(x1,y1),(x2,y2),…,(xn,yn)} where each xiRd was a d-dimensional vector, i.e., each observed data point was composed of the values of d variables, and each yiR was a real value. What we had to do was to find a map f:RdR that minimized the sum of squared errors based on the observed data points. The optimization objective is given as follows:

where β ∈ Rd is the optimized coefficient.

If Eq. 16 is expressed in matrix form, denoted by X = [x1;x2;⋯;xn]T, where each data point xi was regarded as a column vector, then XRn×d, denoted as y = (y1,y2,⋯,yn)T, then the optimization objective in matrix form is given as follows:

Lasso added the L1 regularization term (see Eq. 18) to make the model avoid over-fitting.

Then, the optimization objective function of Lasso is expressed as the following equation:

PLS model was a many-to-many linear regression modeling method, i.e., there are multiple independent variables and multiple dependent variables. It found the best functional fit for a set of data by minimizing the sum of squared errors.

The general multivariate underlying model of PLS is given by the following equations:

where X is a n×m prediction matrix, Y is a n×p response matrix; T and U are n×l matrices and both of them are the projections of X and Y in the higher dimensional space; P and Q are the orthogonal loading matrices of m×l and p×l, respectively, and the matrices E and F are error terms, normally distributed random variables subject to independent and identical distributions. Decompose X and Y to maximize the covariance between T and U.

Gene-Level Analyses

We analyzed the voxel-based features using gene-level analysis. First, quality control (QC) was performed using the PLINK version 1.9 software1 (). We performed genome-wide association studies (GWASs) using the image data and genetic data in whole brain using the linear regression in PLINK. Age, gender, education, and the top 10 principal components from population stratification analysis were included as covariates. A total of 5,574,300 single-nucleotide polymorphisms (SNPs) were obtained by QC. We applied ECS method () to assign SNPs’ to autosomal genes. Then the significant genes was obtained by Bonferroni correction (family-wise error rate p-value < 0.05).

Pathway-Level Analyses

Using the resulting genes, we performed the pathway analysis to assess the biological significance of these features ().KOBAS-I () pathway analysis tool (KOBAS; bioinfo.org) and the Kyoto Encyclopedia of Genes and Genomes database were applied to pathway analysis of the identified genes (P < 0.001).

Results

In recent studies, machine learning was used to detect the subjects and brain regions of AD () and the brain functional statuses of EMCI () and to identify AD and MCI (; ). In this work, we applied a novel feature extraction method and SVM to obtain the features classified EMCI, LMCI, AD, and HC.

Comparison of the Five Methods

We employed the test set to evaluate the classification capability of the five methods, and the experiments were repeated 10 times with the selected parameter combination in each method. As shown in Figure 2, the RS-SVM model has the best prediction accuracy. The AD-HC group had more than 90% prediction accuracy, while the other four methods all peaked below 90%. The prediction accuracy of both the EMCI-HC and LMCI-HC groups exceeded 85%, while the peak values of the other four methods were all below 80%. The curves in Figure 2 also showed that RS-SVM had good stability. In ten replicates, the difference in accuracy was less than 10%. These analyses demonstrated the satisfactory classification ability and stability of the RS-SVM model.

FIGURE 2

Machine learning had been gradually maturing and has been applied to the classification and prediction of AD. We applied the validation set to obtain the optimal parameters and the test set to evaluate the classification capability of the five methods. The evaluation metrics of the five methods implemented in EMCI-HC, LMCI-HC, and AD-HC were shown in Table 2. As shown in Table 2, the RS-SVM has the best accuracy, precision, recall, and F-measure. Only the values of RS-SVM increase with the optimal parameters, and the values of other models are stable. In the AD-HC, EMCI-HC, and LMCI-HC groups, the F-measure of RS-SVM in the validation set were 0.91, 0.86, and 0.85 from high to low. In the AD-HC, EMCI-HC, and LMCI-HC groups, the F-measure of RS-SVM in the test set were 0.93, 0.86, and 0.85 from high to low. This also indicated that the RS-SVM model was scalable, and SVM combined with other schemes have better performance than single SVM. In addition, since the same features were applied to the five models, good results were obtained for all five models (all above 0.8). This proved that the identified features were excellent in the classification of AD and HC and were meaningful for the identification of AD. Therefore, we performed GWAS of these features to analyze their biological significance.

TABLE 2

GroupModelValidation set
Test set
AccuracyPrecisionRecallF-MeasureAccuracyPrecisionRecallF-Measure
EMCI-HCLinear regression0.670.670.670.670.730.730.730.73
Lasso0.790.790.790.790.800.800.800.80
PLS0.80.80.80.80.820.810.810.81
SVM0.730.730.730.730.760.760.760.76
RS-SVM0.860.860.860.860.860.860.860.86
LMCI-HCLinear regression0.620.620.620.620.780.780.770.77
Lasso0.800.800.800.800.810.810.810.81
PLS0.650.640.650.640.660.650.660.65
SVM0.730.730.730.730.740.740.740.74
RS-SVM0.850.850.850.850.850.850.850.85
AD-HCLinear regression0.850.850.840.840.840.850.840.84
Lasso0.850.850.850.850.850.850.850.85
PLS0.910.920.910.910.910.920.910.91
SVM0.870.870.870.870.870.870.870.87
RS-SVM0.910.910.910.910.930.930.930.93

Test results of different models.

Bold fonts represented the model and experimental results in this paper.

Results of Gene-Level Genome-Wide Association Study

We performed the conditional gene-based association scans on whole genome. All genes with conditional association P-values passing Bonferroni correction for family-wise error rate at 0.05 were extracted. We performed the gene-based association analysis by using P-values of 113 novel voxel-based features for identifying susceptibility genes of AD. There are 242 genes (corrected P-value < 0.001) associated with AD. These top 10 conditionally significant genes are shown in Table 3. Studies have shown that CSMD1 (SNP: rs34464519, CorrectedP: 1.74556E-36) was related to AD (; ; ). RBFOX1 (SNP: rs55642412, CorrectedP: 3.18755E-23) has been found to play a role in neuronal development (). PTPRD (SNP: rs62538998, CorrectedP: 1.92988E-19) has been confirmed to be related to AD and MCI in previous studies (). DLGAP2 (SNP: rs72507619, CorrectedP: 3.74049E-17) was found to be predominantly expressed in the brain and associated with a wide variety of neurological disorders (). WWOX gene has been reported to be a potential mechanism that may be involved in the pathogenesis of AD, focusing on the cell death signaling pathway in neurons ().

TABLE 3

No.ChrGeneCorrectedP
18CSMD11.74556E-36
216RBFOX13.18755E-23
316CDH131.07119E-20
49PTPRD1.92988E-19
58DLGAP23.74049E-17
611CNTN54.81385E-16
77MAGI25.93057E-16
820MACROD21.50704E-14
916WWOX1.64798E-14
103CNTN41.87567E-13

Top 10 conditionally significant genes were obtained. Chr represents Chromosome; Gene represents the gene name; CorrectedP represents P-value generated by Bonferroni correction.

Results of Pathway-Level Genome-Wide Association Study

Detecting pathways may provide useful information about the pathogenic molecular mechanism underlying AD. In our work, 70 enriched pathways were identified. The top 10 significant pathways are shown in Table 4. Impaired insulin secretion was associated with higher risk of any dementia and cognitive impairment (). Oxytocin signaling pathway was neuroprotective to many neurological disorders, such as AD (). Vascular smooth muscle contraction was associated with the development of neurodegeneration in AD ().

TABLE 4

NO.PathwaysCorrected P-valueGene
1Insulin secretion1.01E-06PLCB1, PRKCB, PRKCA, CREB5, RYR2, CHRM3, KCNMA1, RAPGEF4, CACNA1C
2Oxytocin signaling pathway4.80E-06PLCB1, PRKAG2, PRKCA, CACNB2, RYR3, RYR2, PRKCB, CACNA1C, ITPR2, CACNA2D3
3Salivary secretion7.70E-06PLCB1, PRKCA, RYR3, PRKCB, CHRM3, KCNMA1, PRKG1, ITPR2
4Vascular smooth muscle contraction7.94E-06PLCB1, CACNA1C, PRKCH, PRKCA, PRKCB, PRKCE, KCNMA1, PRKG1, ITPR2
5Calcium signaling pathway1.48E-05PLCB1, PRKCB, ERBB4, PRKCA, RYR3, RYR2, CHRM3, CACNA1C, ITPR2, PDE1A
6Glutamatergic synapse2.08E-05PLCB1, CACNA1C, PRKCA, GRIK2, PRKCB, DLGAP1, ITPR2, GRM7
7Morphine addiction5.00E-05PRKCA, PDE1A, PRKCB, PDE3A, GABRB3, PDE4D, PDE10A
8Circadian entrainment5.56E-05PLCB1, PRKCB, PRKCA, RYR3, RYR2, CACNA1C, PRKG1
9Pancreatic secretion5.56E-05PLCB1, PRKCB, PRKCA, RYR2, CHRM3, KCNMA1, ITPR2
10Aldosterone synthesis and secretion5.56E-05PLCB1, CACNA1C, PRKCA, CREB5, PRKCB, PRKCE, ITPR2

Top 10 significant pathways.

Discussion

We proposed a voxel-based three-level analysis framework for AD that extracted the voxel-based ROI, which included the whole brain from MRI. Although the voxel-based research could solve the limitations of the research method based on the ROI, it was more easily affected by the characteristics of high-dimensional data. Feature selection in RS-SVM model solved the dimensional disaster caused by too many attributes. In this work, we identified 136, 141, and 113 MRI features for EMCI-HC, LMCI-HC, and AD-HC groups, respectively.

We performed RS-SVM model to identify important brain regions such as hippocampus, amygdala, angular gyrus, and calcarine sulcus for AD. The hippocampus was located in the midlimbic system of the brain and had an important impact on memory and cognitive function. Many studies had shown that abnormalities in hippocampal volume and function were closely linked to AD. Although many patients had not shown symptoms of AD during MCI period, the temporal lobe located in the inner part of the brain had obvious symptoms. The atrophy of the hippocampus was the most obvious (; ; ). Located at the bottom of the brain, the amygdala was shaped like an almond and was the center of the brain to control and manage emotions. Pathological protein changed in the amygdala affected the occurrence, development, and evolution of AD and led to nerve damage or tissue cell aging (; ). The angular gyrus was the portion surrounding the end of the superior temporal sulcus in the temporal lobe. Gray matter atrophy gradually spread from the basal ganglia to the angular gyrus, temporal regions, and eventually to the subcortical-cortical network as neurological disease progresses. The angular gyrus was key AD-risk region (; ). The calcarine sulcus was located posterior to the medial surface of the hemisphere. Less activation of the bilateral anterior calcarine sulcus was associated with better delayed recall in amnestic MCI patients (; ).

In our work, the EMCI-HC, LMCI-HC, and AD-HC groups were classified in order to increase the specificity and detail of classification and to enable early diagnosis of the disease. In this study, five classification models of linear regression, Lasso, PLS, SVM, and RS-SVM were applied. The prediction accuracy of AD-HC was the highest. EMCI and LMCI represented the middle stage of disease progression. Therefore, the prediction accuracies of EMCI-HC and LMCI-HC were lower than that of AD-HC.

The machine learning was applied to classify AD and HC in previous studies (; ; ; ; ). For further classification validation, we plot the ROC curves (Figure 3) of five machine learning classification methods for EMCI-HC, LMCI-HC, and AD-HC groups. The AUCs of the RS-SVM for EMCI-HC, LMCI-HC, and AD-HC were 0.898, 0.839, and 0.964, respectively. As can be seen in Table 2, the proposed classification method, based on RS-SVM, consistently outperforms other methods (i.e., linear regression, Lasso, PLS, and SVM). This proved that the framework based on RS-SVM was optimal compared with the other four methods.

FIGURE 3

With the replication of statistical gene-level GWAS, we obtained PLCB1, PTPRD, RYR3, CACNA1C, KCNMA1, and PRKCA genes. The PLCB1 gene was implicated in AD pathogenesis (). The PTPRD gene has been found on the short arm of human chromosome 9. Recent report identified PTPRD association with the extent of neurofibrillary pathology in AD brain specimens (; ), and also RYR3 gene for the RYR which functions to release the stored endoplasmic reticulum calcium ions (Ca2+) to increase intracellular Ca2+ concentration. The studies demonstrate that altered levels of intracellular Ca2+ affect neurodegeneration (; ). It has been shown that the expression of CACNA1C inhibits the hyperphosphorylation of Tau protein ().

Some pathways include insulin secretion, oxytocin signaling pathway, salivary secretion, vascular smooth muscle contraction, and AD closely related to genes (Figure 4).

FIGURE 4

In summary, our proposed framework based on RS-SVM performed well in features constructed, and the framework had good classification performance for EMCI-HC, LMCI-HC, and AD-HC groups. In particular, AD-HC was the best in terms of classification accuracy. Several pathogenic genes and abnormal subregions identified singing this framework are related to AD. Therefore, we speculate that the remaining genes identified could be regarded as the candidate genes for AD. The discoveries in this study provide new candidate genes for AD, and the constructed features can be regarded as a new indicator to distinguish AD from HC.

Publisher’s Note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Statements

Data availability statement

Rs-fMRI data were downloaded from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database (http://adni.loni.usc.edu/). Application for access to the ADNI data can be submitted at http://adni.loni.usc.edu/data-samples/access-data/.

Ethics statement

The studies involving human participants were reviewed and approved by Data collection and sharing for this project was funded by the Alzheimer’s Disease Neuroimaging Initiative (ADNI). We applied the acess from ADNI. The patients/participants provided their written informed consent to participate in this study. Written informed consent was obtained from the individual(s) for the publication of any potentially identifiable images or data included in this article.

Author contributions

XM, YWu, WL, and ZJ led and supervised research. XM, YWu, and WL designed the research and wrote the article. XM, YWu, and ZX performed features extraction and selection and random survey support vector machines. WL, YWa, and ZX performed data preprocessing and quality control. XM and WL did gene- and pathway-level analysis. All authors reviewed, commented, edited, and approved the manuscript.

Funding

This research was funded by the National Natural Science Foundation of China (61901063 and 51877013), MOE (Ministry of Education in China) Project of Humanities and Social Sciences (19YJCZH120 and 21YJAZH091), National Statistical Science Research Project (2020LY074), and the Science and Technology Plan Project of Changzhou (CE20205042 and CJ20210155), Jiangsu Provincial Key Research and Development Program (BE2021636), and Key Lab of Intelligent Optimization and Information Processing Minnan Normal University (ZNYH202007). This work was also sponsored by Qing Lan Project of Jiangsu Province. Data collection and sharing for this project was funded by the Alzheimer’s Disease Neuroimaging Initiative (ADNI) (National Institutes of Health Grant U01 AG024904) and DOD ADNI (Department of Defense award number W81XWH-12-2-0012).

Acknowledgments

ADNI is funded by the National Institute on Aging, the National Institute of Biomedical Imaging and Bioengineering, and through generous contributions from the following: AbbVie, Alzheimer’s Association; Alzheimer’s Drug Discovery Foundation; Araclon Biotech; BioClinica, Inc.; Biogen; Bristol-Myers Squibb Company; CereSpir, Inc.; Cogstate; Eisai Inc.; Elan Pharmaceuticals, Inc.; Eli Lilly and Company; EuroImmun; F. Hoffmann-La Roche Ltd and its affiliated company Genentech, Inc.; Fujirebio; GE Healthcare; IXICO Ltd.; Janssen Alzheimer Immunotherapy Research & Development, LLC.; Johnson & Johnson Pharmaceutical Research & Development LLC.; Lumosity; Lundbeck; Merck & Co., Inc.; Meso Scale Diagnostics, LLC.; NeuroRx Research; Neurotrack Technologies; Novartis Pharmaceuticals Corporation; Pfizer Inc.; Piramal Imaging; Servier; Takeda Pharmaceutical Company; and Transition Therapeutics. The Canadian Institutes of Health Research is providing funds to support ADNI clinical sites in Canada. Private sector contributions are facilitated by the Foundation for the National Institutes of Health (www.fnih.org). The grantee organization is the Northern California Institute for Research and Education, and the study is coordinated by the Alzheimer’s Therapeutic Research Institute at the University of Southern California. ADNI data are disseminated by the Laboratory for Neuro Imaging at the University of Southern California.

Conflict of interest

The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

References

Summary

Keywords

Alzheimer’s disease, RS-SVM, voxel-based features, gene-level, pathway-level

Citation

Meng X, Wu Y, Liu W, Wang Y, Xu Z and Jiao Z (2022) Research on Voxel-Based Features Detection and Analysis of Alzheimer’s Disease Using Random Survey Support Vector Machine. Front. Neuroinform. 16:856295. doi: 10.3389/fninf.2022.856295

Received

17 January 2022

Accepted

08 February 2022

Published

28 March 2022

Volume

16 - 2022

Edited by

Yu-Dong Zhang, University of Leicester, United Kingdom

Reviewed by

Xia-an Bi, Hunan Normal University, China; Bingrong Xu, Wuhan University of Technology, China; Simin Li, Southern Medical University, China

Updates

Copyright

*Correspondence: Zhuqing Jiao,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics