ORIGINAL RESEARCH article

Front. Earth Sci., 10 July 2026

Sec. Geoinformatics

Volume 14 - 2026 | https://doi.org/10.3389/feart.2026.1748893

A framework for earthquake seismic signal classification with the aid of hybrid feature selection schemes and robust machine learning models

  • 1. Department of Artificial Intelligence Convergence, College of Information Science, Hallym University, Chuncheon, Republic of Korea

  • 2. College of Medicine, Hallym University, Chuncheon, Republic of Korea

  • 3. Department of Population and Quantitative Health Sciences, University of Massachusetts Chan Medical School, Worcester, MA, United States

Abstract

One of the most catastrophic natural events is an earthquake, which causes huge destruction to humanity, thus necessitating the use of seismic signals to detect and classify them. Many machine learning models have been proposed that act as automated earthquake detection and classification methods by using seismic signals. However, the proposed method in this study utilizes the concept of learning-based feature extraction from seismic signals such as graph-embedding learning (GEL), through low-rank representation (LRR) and robust neighborhood embedding (RNE). In a parallel approach, statistical feature extraction approaches such as cross-correlation function (CCF) and partial CCF (PCCF) are implemented with multivariate variational mode decomposition (MVMD). Kriging model-based feature selection and K-best features selection techniques are then utilized, and the features are finally classified with the standard machine learning classifiers. This combined approach is considered a novel technique. The second technique proposed in this study utilizes the concept of reinforcement learning (RL) Bayesian networks (BNs) with the AdaBoost-based Elman neural network (ENN) model known altogether as the “RL–BN–AENN model” with extracted and selected features. Third, a differential evolution (DE)-based Bayesian classifier is proposed, the DE-BC model, with extracted and selected features. Finally, the concept of linear discriminant analysis (LDA) is utilized with a fuzzy weighted relevance vector machine, the LDA–FW–relevance vector machine (RVM) model, and the same LDA is utilized with a fruit fly optimization algorithm (FFOA) and support vector machine (SVM) concept, the LDA–FFOA–SVM model, which is utilized with extracted and selected features. The results show that a high classification accuracy of 98.89% is obtained for the two-class classification problem when LLR with Kriging model feature selection and the AdaBoost–ENN classifier are implemented. A high classification accuracy of 92.95% is obtained for the three-class classification problem when PCCF with MVMD and Kriging model feature selection classified with the LDA–FFOA–SVM classifier is implemented.

1 Introduction

Due to unforeseen liberation of energy in the lithosphere, the surface of the earth is easily shaken, creating seismic movements that results in earthquakes or tremors (). Earthquakes can be assessed at different intensity levels, with some tremors having no impact while others causing very great damage to buildings, nature, animals, and human lives (Schweier and Markus, 2006). The type and frequency of an earthquake depend on seismic activity occurring at a particular time. Some earthquakes happen naturally, while others occur due to human activity such as nuclear weapons testing and mining. Earthquakes are often attributed to geological faults, while other factors such as landslides and volcanism also contribute to earthquakes (). Soil liquefaction is a major effect of earthquakes, which often leads to great loss of life. Tectonic movements influence the occurrence of earthquakes; the governance of rupture dynamics is investigated using elastic-rebound theory (). The management of earthquake risks includes earthquake engineering and planning and involves features like preparedness, forecasting, and prediction. Many religious beliefs and myths are associated with the cultural influence of earthquakes, and these have a profound effect on human societies (). When a large amount of elastic strain energy causes fracture propagation along the land surface, a tectonic earthquake occurs. Along the surfaces of a fault, if there is no irregularity, then the fault sides move aseismically, increasing frictional resistance and leading to stick–slip behavior (). When stress then increases high enough to break the asperity, the stored energy is released over the closed side of the fault. The combination of strain seismic waves and the cracking of rock results in an earthquake. The three types of faults are normal, reverse, and strike–slip faults (). The size of an earthquake can be assessed using the Richter instrumental scale. The amplitude of an earthquake can be assessed through this simple measurement, and its utility has become minimal nowadays. Seismometers are used to record the seismic waves that propagate through the interior layers of the Earth (). To assess remote earthquakes, the surface-wave magnitude was implemented so that accuracy could be improved for large events. The intensity and magnitude of an earthquake are not directly related and are computed using different techniques. Various types of seismic waves are produced by an earthquake, and these travel with varying velocities such as longitudinal P waves, transverse S waves, and surface waves (). The velocity and speed of seismic waves depend on two different factors: the elasticity and the density of the propagation medium. Shaking and ground rupture is the main problem caused by earthquakes, resulting in damage to large buildings and rigid structures. The disruptive side effects of earthquakes include severe disruption to public transportation networks, the interruption of water, power, and gas connections, loss of communication systems, and general property damage (). The aftermath of an earthquake causes a range of psychological effects such as anxiety, depression, and panic attacks to survivors. Depending on the damage levels and the financial status of the affected community, recovery time may vary from a few months to a few years. Finally, earthquakes are sometimes accompanied by other natural disasters such as landslides, fires, tsunamis and floods, thereby aggravating the situation (). Therefore, it is useful to analyze earthquake classification and detection using seismic signals as this would be highly beneficial for the betterment of humanity. A large body of research has been reported regarding earthquake classification using seismic signals and machine learning, with some being reported in this study.

A seismic arrival time picking technique has been developed with a concept called “PhaseNet”; it uses convolutional neural networks (CNNs), with the detection of P and S wave arrivals reported by Zhu and Beroza (2019). A novel feature extraction technique with a supervised machine learning technique was proposed by where the classification of noise versus earthquake was attained in a static and dynamic environment (). The raw waveform data and CNN have been used for the rapid prediction of earthquake ground-shaking intensity, intensity measurements being accurately predicted by . The capsule neural network (CapNN) was used for to estimate P-waves by Saad and Chen (2020), such that the detection of earthquakes was perfect. An early earthquake warning detection system using the Internet of Things and a multi-head CNN for the classification of earthquakes was reported by Wu et al. (2021). With the help of smart devices, a collaborative approach was developed for the detection of earthquakes using artificial neural networks (ANNs) (). A federated-learning-concept-based machine learning concept was implemented for earthquake prediction by Tehseen et al. (2021). A pipeline-based CNN with the aid of single-station waveforms was implemented by ( for-earthquake detection and classification. For the exclusive purpose of seismic phase classification, CapNN was used by Saad and Chen (2021), which involved the detection of P and S waves in which arrivals were accurately detected. A novel deep-learning approach based on a recurrent neural network (RNN) and 3D CNN was developed by Shakeel et al. (2021) in which earthquake-versus-non-earthquake classification was executed. A long-short term memory (LSTM) with Keras optimizer was used by () to classify earthquakes. For the arrival time estimation of P-waves, a feed-forward neural network was developed by . A batch normalization-based graph CNN was used to predict the magnitude, location, and depth of an earthquake by . Random forest (RF) algorithms were used for the multi-classification detection of earthquakes by . P and S wave arrival detection was conducted using deep learning with the help of a CNN by Zhu et al. (2022). In a seismic waveform, real-time P-wave detection was performed and implemented on a low-maintenance MEMS accelerometer using a CNN by . An attention-based fully convolutional DenseNet was utilized for earthquake detection by . For seismic phase selection and earthquake detection, an attention-based U-shaped neural network was implemented by . For the estimation of magnitude and identification of earthquake, a convolutional recurrent model was developed by . Arrival time detection of P and S waves based on object detection was classified with a CNN by . An adaptive neuro-fuzzy inference system (ANFIS) was used for real-time earthquake prediction by . A novel butterfly pattern using seismic signals and an automated earthquake detection system was developed by . Finally, from a lateral point of view, other factors such as the development, challenges, and potentials of underground energy storage systems for analyzing earthquakes are studied in this experiment (Wang et al., 2026). The thermomechanical behavior and damage mechanism of the lining backfill body of high-temperature thermal energy storage reservoirs in mines (Xu et al., 2026a), along with the thermo-mechanical damage characteristics and fracture behavior of backfill bodies in thermal energy storage cavities (Xu et al., 2026b), are also studied here in detail before proceeding with the experimental analysis.

The experimental workflow is implemented in this study is as follows:

  • Four different strategies are proposed. Once the basic pre-processing of the seismic signals is explained using simple independent component analysis (ICA), the main research is proposed. This proposed method utilizes the concept of learning-based feature extraction from seismic signals such as graph embedding learning (GEL), learning through low-rank representation (LRR) and robust neighborhood embedding (RNE). In a parallel approach, a statistical feature extraction approach, including cross-correlation function (CCF) and partial CCF (PCCF), is implemented with multivariate variational mode decomposition (MVMD), and then Kriging model-based feature selection and K-best features selection techniques are utilized, followed by classification with the standard machine learning classifiers.

  • The study then utilizes the concept of reinforcement-learning-based Bayesian networks with an AdaBoost-based Elman neural network (ENN) model—the RL–BN–AENN model—with the extracted and selected features.

  • Next, a differential evolution (DE)-based Bayesian classifier (BC) is proposed—the DE–BC model—with the extracted and selected features.

  • Finally, the concept of linear discriminant analysis (LDA) is utilized with a fuzzy weighted relevance vector machine—the LDA–FW–RVM model—and the same LDA is utilized with the fruit fly optimization algorithm (FFOA) and a support vector machine (SVM), a concept collectively known as the “LDA–FFOA–SVM model”, which is then utilized with the extracted and selected features.

The remainder of this study is organized as follows. Section 2 discusses the feature extraction and selection schemes for seismic signals. Section 3 explains the RL–BN–AENN classification model. Section 4 explains the DE–BC classification model, and Section 5 explains the proposed hybrid LDA–FWRVM–FFOA–SVM classification model. The results and discussion are explained in Section 6, and the study concludes with Section 7.

2 Feature extraction and feature selection schemes for seismic signals

To the best of our knowledge, the following schemes were the first of their kind implemented for the feature extraction and selection of seismic signals.

2.1 Learning-based feature extraction for seismic signals

The proposed method utilizes the concept of learning-based feature extraction from seismic signals, such as GEL, LRR, and RNE.

2.1.1 Graph-embedding learning

For classification, a vital role is played by the local manifold composition of the data. The original and actual neighbor relationships can be preserved very well by promoting the discriminating power in many learning-dependent techniques. Basic geometric intuition is required for the assessment of GEL (). For a particular seismic data matrix , it is generally presumped that every data point along with its corresponding neighbor is placed on local linear planes of a particular manifold. The linear coefficients are utilized to assess these planes so that each datum can be reconstructed well from its neighbors. With the help of KNN techniques, adjacent graphs can be constructed to make the description of the geometry structure easier (Taunk et al., 2019). From node to node , if an edge is present, then the weight is indicated as . By solving the optimization problem, the solution for the weight matrix is traced and is represented as follows:where the neighbor set of is represented as and the reconstruction coefficient is represented as . As the objective function is subjected to reconstruction errors, neighborhood preserving embedding (NPE) arose to help preserve the identity of original data by holding the neighborhood composition of the actual data in a low-dimensional space. The original data are projected to a line so that they can be specified as a linear combination of the original neighbors with the respective coefficients . If a linear projection is assured, it can be specified as , where , .

By means of solving the optimization problem, the generation of the projection matrix is performed:

With the aid of minimum eigen values, the optimal projection matrix can be elucidated so that, for the given data, the embedding vector in the low-dimensional space can be easily identified.

2.1.2 Learning through low-rank representation

The specification of seismic data can be achieved as a linear combination of bases in the dictionary , that is , where specifies a representative coefficient matrix. In the dictionary in every subspace, the data matrix is utilized; therefore, every data point is self-represented as a simple linear combination of multiple data points (Wang et al., 2024). The contribution of every datum can be characterized by the coefficient matrix so that the reconstruction of data can be performed well. In learning through the low-rank representation model, the sparsity of the singular values is assessed so that the low-rank representative matrix can be clearly analyzed. For the representation matrix , the nuclear norm is identified and is utilized as a surrogate so that global information can be captured effectively. The uncovering of the correlation structure of data can be carried out well and, thus, by means of solving the optimization issue, the optimal coefficient matrices are identified as follows:where the nuclear norm is represented as and the regularization parameter is represented as . is used to model corruption in the data, and optimization is usually carried out using the Lagrange multiplier technique. Once is obtained, the affinity graph is determined by means of a construction of .

2.1.3 Robust neighborhood embedding

Based on the assessment of the weight matrix and indicator matrix of seismic data, RNE is utilized for unsupervised feature selection where the indicator matrix is expressed as follows:where the feature index subset is represented as —that is, . The indicator matrix is column orthogonal. Based on the indicator matrix , the assessment of the index subset of chosen features can be determined. For data , new data is obtained by choosing features from its original set . A similar local neighborhood relationship is present for both the original data and new data. The formulation of the RNE model is as follows:

Therefore, many meaningful features are chosen from the original data through the index matrix, and the local neighborhood relationship is preserved well by the new data with respect to the original data (). Equation 7 can be represented as follows:

Therefore, a weight matrix is derived in advance so that the reconstruction errors can be well minimized at the end.

2.2 Statistical feature extraction with MVMD

In a parallel approach, statistical feature extraction methods such as CCF (Yang et al., 2021) and PCCF analysis () are implemented with the MVMD in this study.

2.2.1 CCF and PCCF analysis

Assuming that the seismic data matrix is split into two data variable parameters and with and lag , then between the two-time series and , the covariance function is represented as follows:where the means of and are computed with the help of the covariance function, expressed as follows:where the standard deviation values of and are represented as and , respectively. The correlation between and is established if is a positive integer. The correlation between and is established if is a negative integer. The correlation coefficient is considered when , thus expressed as . When a correlation is present between the signal and a time-delayed signal and a strong correlation is present between signal and , , then the linear influences of the main control signal can be separated easily using PCCF, which is expressed as follows:where the PCF at lag is represented by and . A discrete PCCF is produced, so the linear influences can be eliminated easily (). The range of PCCF is usually between −1 and 1, and the cross-correlation is reflected by the signal residuals and . The generalization of seismic signals is conducted, and the recursive determination of the order PCCF is assessed. Between variables and , the PCCF can be isolated easily, and the order PCCF is expressed as follows:where the input variable is represented by , the output is represented by , and the lag is represented by .

2.2.2 MVMD

An improved version of VMD is multivariate VMD (). Highly complex seismic signals can be decomposed into sub-signals with this non-recursive decomposition scheme with the aid of intrinsic mode function (IMF). Within the seismic signal, the important modes can be isolated adaptively with the help of VMD. A fixed number of significant modes expressed as can be extracted from real signal with channels. The set of individual channels is projected aswhere and the signal is referred to as . A collection of multivariate patterns is isolated from the data. The aggregate bandwidth should be minimized, and the original signal should be reconstructed faithfully. With the help of the Hilbert transform (), the analytic signal which corresponds to is expressed as . By utilizing the L2 norm, the determination of is achieved, and the cost function of MVMD is expressed thus:

For the harmonic mixture of the full vectors, a single frequency component is utilized. Multivariate oscillation is identified so that the frequency components can be identified correctly. The Frobenius norm is computed so that the objective function can be well represented as follows:where indicates the analytical modulated signal. For the MVMD, the optimization problem is expressed as follows:subject to .

By employing the alternating direction techniques of multipliers, the above listed complex optimization issue is resolved, and hence the main intent is to decompose the main issue into sub-issues so that it can be managed easily. Updating the center frequency and mode can easily be achieved in this process.

2.3 Feature selection for seismic signals

A K-best features selection and Kriging model are used to select the best features and, finally, classified with multiple machine learning classifiers. This technique is a novel contribution of this research.

2.3.1 K-best feature choosing model

Many feature selection techniques have been utilized in the past to improve model accuracy and prevent overfitting, thus reducing the training time of the model. One prominent method utilized is the K-best feature selection technique (Zhang et al., 2021). The leading K attributes are chosen depending on a particular statistical value. A univariate feature selection technique is used as a fundamental approach here. From many univariate statistical tests, K features with the highest scores are chosen with the help of the F-test linear relationship between two probabilistic variables. With the help of F-statistics, high F-values are utilized to guide the selection process. To analyze the influence of individual regressors, a rapid linear model is utilized in this technique. The computation of the correlation coefficient is computed between the signal input and output to assess the F-statistics, which is expressed as follows:

The sum of squares based on the concept of regression is indicated by the numerator, and the sum of squared errors is expressed by the denominator. The degree of freedom is expressed by the term . The relative power of regression is examined using F-statistics. The chosen K-features with the highest F-values are expressed here as the ultimate values, as this method seems to be quite versatile and robust.

2.3.2 The Kriging model as a feature selection technique

The Kriging model is often considered a surrogate model (W et al., 2022). It is not a traditional feature selection algorithm, but it acts as an embedded or filter method in spatial analysis. The best spatial features can be easily selected with the help of the Kriging model, and it can handle spatial dependency very well and so is used in this study. On estimated values, a range of uncertain information can be processed using this surrogate technique. It mainly employs Bayesian optimization and is thus highly dependent on statistical learning and is more accessible to kernel techniques. Uncertainty information is used to update the Kriging model. For an individual , the objective function value is approximated as follows:where a polynomial function of is represented by , which represents the regression model. The Gaussian distribution is expressed as and is expressed as follows:where the variance is expressed as and the mean is represented as . The sampled data points are interpolated using the Kriging model and are represented as . The covariance matrix is used to assess the local deviators and is expressed as follows:where the correlation matrix is represented by and is an matrix. The correlation function is represented by . Kriging model training is conducted with the help of training samples. An estimated objective value is always identified for every input. When using the Kriging model, the computational complexity can be slightly high and can still deal with a high number of input parameters (W et al., 2022). In the design space, the sample point location is used to assess the productivity of the Kriging model. When the sample points are far away, the Kriging model works efficiently.

2.4 Machine learning classifiers utilized

The classifiers utilized in this study are commonly used classifiers such as AdaBoost, naïve Bayesian classifier (NBC), random forest (RF), decision trees (DTs), and logistic regression (LR). Figure 1 shows the proposed block diagram of the feature extraction and selection module with the machine learning classifiers for earthquake classification using seismic signals.

FIGURE 1

3 RL–BN–AENN (reinforcement learning Bayesian networks with AdaBoost-based ENN model)

3.1 Reinforcement learning

A sequential decision-making issue is solved by RL, where the policy is optimized by an RL agent by means of active interaction with the environment (Qiang and Zhongli, 2011). A Markov decision process (MDP) is used to largely influence this effect. A standard MDP is defined by a set . The state space is denoted as with size and action space is denoted as with action space . The reward distribution function is denoted by with . This indicates that a time step , the instantaneous reward for considering action in state is determined. The transition probability is denoted by , with specifying the probability of transitioning from to upon action . The discount factor is indicated by . An action is chosen by the RL agent based on for a given policy . It is then transited to the preceding state , thus ensuing , and an instant reward is received. With the help of a discounted factor , the total discounted reward at is defined by

The optimal policy is attained by the RL agent object to maximize the expectation of . The parameter is used to parametrize the policy network and assess the policy . The value function is used and expressed as

The definition of the action value function is determined as

The expectation operator with respect to is denoted as

Through the gradient ascent technique, the policy parameter is updated out using the policy-based RL technique:where the learning rate is assessed by and is the total expected reward, which is approximated as follows:

The advantage function is denoted as and is obtained by subtracting by . It is represented as follows:where after the consideration of and , the expected reward that the agent could receive is expressed by .

3.2 Bayesian networks

A Bayesian network comes under the category of probabilistic graphical models (PGMs); it is defined as a set , where is the directed acyclic graph (). Here, the set of nodes is specified by , and the edges are represented by . is the probability function. The classification of the Bayesian or discrete Bayesian networks is based on whether the variables are continuous or discrete in nature. To deal with the incessant attributes, the Rayleigh distribution can be utilized in addition to the Gaussian distribution. In this study, only discrete Bayesian networks are utilized. Initially the structure and probability function of the Bayesian network are obtained. can be parametrized by a conditional probability table. The topology is structured dependent on the causality of nodes. Between the states and action, the causal relationship can be easily expressed by the Bayesian network structure, as the state input and action output are known for most RL tasks. The favorable parameter of the probability function is estimated along with the estimation of probabilistic inference of Bayesian networks. Figure 2 shows the RL–BN–AENN model.

FIGURE 2

3.3 Parameter estimation

The probability function of all the nodes is learned thoroughly, where every node in the Bayesian network indicates a variable. From the data, the conditional independence of all the nodes can be easily learned from the structural entity of the Bayesian network. For dataset comprised of the entire observed samples of a Bayesian network and for the parameter estimation technique, maximum likelihood estimation (MLE) is utilized (). It is assumed that a Bayesian network has nodes , and its probability function is parameterized by . It is also presumed that for the node in , it has candidate values, and its present nodes have candidate combinations. Every parameter that indicates the conditional probability between node and its present node when and parent is expressed as , where . Based on the probability property, the accumulated quality of the candidate value of is represented as follows:

The optimal parameter is learned using the MLE technique by enhancing the favorable likelihood between the parameter and the dataset, and it is expressed as follows:where the likelihood function of is expressed as . The number of samples that satisfy parent is .

The optimal is obtained using the Lagrange operation as follows:

MLE is used for parameter estimation in this study.

3.4 Probabilistic inference

Probability inference of the Bayesian network is used to assess the posterior probability of target variables (). The probability distribution of variables can be easily computed using this method. Computational efficiency is also improved using such techniques. The joint probability distribution for a Bayesian network is expressed as follows:where the probability destiny function is denoted by ,

Only a simple Bayesian network is used here, so the exact inference methods can be easily applied to it. Variable elimination is used to decompose the joint probability distribution function; thus, a succession of conditional probability outcomes represents a Bayesian network. To obtain the marginal probability , the variables are eliminated using the VE method as follows:

3.5 AdaBoost–ENN hybrid model

The ENN is a recurrent neural network with feedback connections and local memory units (). The specialty of the ENN is that along with input layers, output layers, and hidden layers, there is also a context layer. The main purpose of the latter is that it acts as a one-step delay operator so that memory is achieved. Additionally, it has the notable ability of adapting itself based on time-varying characteristics. The effectiveness of the dynamic process system can be directly reflected by it. The ENN represents a significant improvement compared to the forward neural network, but a drawback of it is that the recursive nature of the hidden layer cannot be modified because it is fixed. To address this problem, the AdaBoost algorithm is used to enhance the ANN and increase prediction performance (Zerrouki et al., 2017). A beautiful example of a boosting algorithm that regresses with time series data is the AdaBoost algorithm. By repeatedly training the same dataset repeatedly, various weak learners can be generated by the AdaBoost algorithm, and these weak learners can then be hybridized into powerful learners to achieve a high generalization effect. The weights of the samples are increased by the AdaBoost algorithm with high prediction error. The mitigation of weak learners with poor learning abilities is implemented, and the weights of samples with a high training effect are obtained.

The AdaBoost–ENN model is explained as follows:

  • The original data are pre-processed via quantification and normalization.

  • The training set is assumed to be , .

  • In the training set, the initialization of the first distribution weight of the sample is determined to be .

  • Assessment of the neural network structure is conducted based on the input and output dimensions, and the initialization of the weights and thresholds of the ENN is then performed.

  • The weak prediction is identified.

  • When training of the weak prediction is carried out, the training of the ENN is also carried out with the training set, and the output is given as prediction results. For the prediction series , the sum of the prediction error is obtained; it is expressed as follows:where is the prediction result and is the original value.

  • The weights are updated and, based on , computation of the weight of the series is conducted thus:

  • Weight adjustment of the next training sample is performed and is expressed aswhere the normalization font is expressed as and

  • With the help of T-round training, the T strong prediction functions are obtained as

  • The formation of the strong prediction function is expressed as

3.6 Weight assessment method

The weight computation process is as follows:

  • The number of feature sets is assumed to be . For each feature set, prediction accuracy is defined as and expressed as follows:where indicates the precision quality of the samples in the feature set, is the predicted value of the sample in the feature set, indicates the original value of the sample in the feature set, and is the average of the actual values in the feature set.

  • For every feature set, the output weight is defined as and is computed as follows:

  • Based on the weights of each feature set, the final prediction results are obtained aswhere the predicted value of the sample in the feature set is expressed as . Based on the weighted technique, is the predicted output.

4 DE-based Bayesian classifier

4.1 DE

By maintaining a population of candidate solutions, a problem can be easily optimized using the DE algorithm (Price, 1996). By hybridizing existing ones, new candidates can be created based on their respective formulas. The candidate solution with the highest score and best fitness is retained. With a population of individuals, the DE algorithm is initiated as . Unless the stopping criterion is satisfied, the three main operations of DE are repeated.

4.1.1 Mutation operation

A donor vector is created by the mutation operator with respect to every population member in the present generation. Some of the mutation strategies used are as follows:

The indices indicate the random and mutually different integers within the range . The mutation scaling factor is denoted by and is within the range of [0.2]. The best individual vector is denoted as .

4.1.2 Crossover

Once the mutation is attained, crossover is performed between the target vector and its respective mutant vector to form the trial vector .

For every of the variables,where the crossover control parameters are represented as CR and are within the range [0,1] and a uniformly distributed random number is represented as . A randomly chosen index is , which ensures that receives at least one component from .

4.1.3 Selection

Once the trial individual is reproduced, it is compared by the selection operator with its respective target individual to determine which one—the target or trial individual—progresses to the next-generation . It is expressed as follows:where the objective function to be optimized is expressed as . This ensures that the next-generation member is quite a fit individual. The target individual is replaced if is better in the next-generation target individual . DE algorithms are continuously developed to improve optimization performance.

4.2 Methodology

The ensemble method based on weighted voting and DE is expressed below. The base classifiers are selected randomly. For all these, the proper weights are found based on prediction confidence through the DE technique. Figure 3 shows the overall proposed feature extraction and selection process with the DE–BC model. Figure 4 shows the proposed DE–BC model alone.

FIGURE 3

FIGURE 4

4.3 Bayesian classifier: selection and training

In the ML community, the advantage of using ensemble classifiers is widely recognized as having good accuracy (Sefat et al., 2021). The selected individual classifiers differ from one another. By manipulating the training examples, diversity is ensured to generate multiple hypotheses. Five base classifiers are here: NB, KNN, AdaBoost, DT, and Bayes Net.

4.4 Parameter selection using the DE model

The weight of every base classifier is considered as an ensemble for parameter optimization. Various parameter settings have various influences on the performance of the model. DE is used to trace for the optimal weights. A random initial population of solution candidates is possessed using DE, and it is enhanced using evolution operators. The maximum iterations are predefined to assess the stopping criterion of DE. The other vital parameters for DE are population size mutation scaling factor , and cross rate . The algorithm for parameter selection using the DE model is given in Algorithm 1.

Algorithm 1

  • Input: Control parameters: Population size , mutation factor , crossover rate .

  • Initialize:

  • Generate uniformly distributed random population with individuals , where is a vector indicating the weights of base classifiers.

  • Assign generation iteration

  • While the finishing criterion is not satisfied do

  • For do

  • Choose random indices to be varying from each other.

  • Calculate a mutant vector

  • Generate random number

  • For do

  • Choose trial individual using (47)

  • End for

  • Compute the fitness of the vector and

  • Employ ten-fold cross-validation model.

  • Update vector of the superseding generation using (48)

  • End for

  • Update generation iteration

  • End while

  • Output: The output weights

4.5 Initialization

A population with individuals is initialized . As a D-dimensional vector, an individual can be defined as , which indicates the weight of base classifiers and the size of base classifiers as represented by . By means of uniform distribution, each individual is generated in the range of [0,1].

4.6 Evaluation of fitness

The DE based classification model is trained using each individual vector, and the fitness function evaluation is performed based on 10-fold cross validation accuracy. Assuming the number of categories and base classifiers to vote, for each sample , the prediction category of the entire weighted voting is expressed aswhere the binary variable is denoted by . If the base classifier matches and classifies the sample into the category, then ; otherwise, the value of . In an ensemble, the weight of the base classifier is represented by , which is optimized and made favorable using the DE algorithm (Algorithm 1).

Accuracy is expressed as follows:

Using DE, the best individuals are found, and the optimal weights are traced so that an ensemble classifier can be generated to classify the datasets.

5 Proposed hybrid LDA–FWRVM–FFOA–SVM model

5.1 LDA

A challenging issue in machine learning tasks is attributed to the high number of features. To achieve high classification accuracy, dimension reduction is utilized. Specifying the high dimensional features in low dimension space is the main intention of dimensional reduction; it is performed to mitigate complexity and reduce the number of features. LDA is used as the dimension reduction technique in this study, and the linear features are easily detected by it, so between-class sample separation is maximized and within-class distribution is minimized (). A training set with samples is considered, and every sample is specified as a column vector of length . The training samples belong to one of the classes in . Assume is the list of samples in class , with indicating the total number of samples in a class. In LDA, the within class distribution matrix is specified as follows:

The between-class matrix is specified by

The mean of the dataset is represented by and is expressed by . The mean of the class is expressed as and is given as . The between-class variance with respect to within-class variance is maximized by the linear transformation , where is represented as the matrix, in which represents a desired number of dimensions. The Eigen vectors are nothing, but the columns of the optimal estimation of and the condition is represented as

This indicates the highest Eigen value of , with indicating the Eigen value. The distribution matrices and are simultaneously transversed by this common outcome, underlining that the data relationship between the within- and between-class is disconnected. Therefore, the low variance columns are removed, and the dimension of the LDA is reduced. Figure 5 shows the proposed LDA–FWRVM–FFOA–SVM model.

FIGURE 5

5.2 Fuzzy weighted relevance vector machine classifier

One of the standard machine learning techniques is RVM, which is dependent on both the Bayesian learning concept and the respective kernel functions. Here, there is no extensive adjustment of the parameters, so it is used for quality estimation. Probability prediction can be enhanced by using adaptive solutions of sparse models. In fuzzy weighted relevance vector machine classifier (FWRVM), for the optimal selection of the weight vectors, fuzzy membership vectors are used (). The assignment of the weights is performed using the features obtained from LDA, thus improving the classification of FWRVM. The posterior probability of each seismic sample is predicted by the FWRVM model. The training set is considered, where samples belong to classes . FWRVM is used to carry out statistical analysis, and by introducing the logistic sigmoid function , the comprehensive linear model is utilized with the estimator function . For the probability , where . The adaptable parameters of FWRVM are indicated by . The radial basis function is denoted by . Due to its efficiency, the RBF kernel function is utilized. Then, the fuzzy membership vectors for the seismic signal samples are introduced, and so is thus calculated as follows:where . The optimal weight vector must be identified once the fuzzy membership vector is adopted. To find the optimal weights, the value of must be determined to maximize the likelihood as follows:

The vector of hyperparameters is denoted by . Analytical assessment of the weights is not possible, and so a marginal probability or weight position or probability is neglected. The weights are approximated using a Laplacian technique. The estimation of the best feasible weights for the current stable value of is determined utilizing the posterior approach position. As is feasible, weight maximization is expressed as follows:where and . Depending on the Gaussian approximation and the log posterior of the weights , the Laplacian approximation is built as follows:

The diagonal matrix is represented as and .

The policy matrix is determined by . The formulation of the covariance matrix can be obtained via the negation and inversion of Equation 59. With the aid of an iterative re-approximation function, updating the hyperparameter can be achieved. The term , which is randomly obtained, is estimated using , where indicates the diagonal ingredient of the matrix. The reapproximation of is expressed as follows:where . The assignment of is carried out, and then and are reapproximated unless satisfactory convergence is obtained.

5.3 SVM

SVM was developed based on statistical learning theory (Wang, 2022). It is used for nonlinear high-dimensional problems, as it is dependent on the structural risk minimization principle. SVM also has a very strong generalizability and deals with the separation of a hyperplane in the feature space which has a high dimension utilized for sample classification. The samples are considered here as is the feature set of seismic signals, is the sample label, and the number of samples is represented by . The optimal separating hyperplane is indicated as . The sample point for separating hyperplanes is nothing but the functional margin and is represented as follows:

Through the normalization of , the geometrical margin is obtained and is expressed as . The maximum value of is obtained and is likely equal to the minimum value of . The issue is transformed into a quadratic programming issue in which the slack variable is written as , and the penalty coefficient is expressed as . For Equation 62, Lagrange duality translation is performed to translate it into a dual problem, as shown in Equation 64:

The optimal separating function is expressed as follows:

The nonlinear problem in the low-dimensional space is translated into a linear problem in high-dimensional space by using the kernel function, and the optimal separating function is expressed as

Kernel functions such as linear, polynomial, RBF kernel, and sigmoid kernel are used. In our study, the RBF kernel is used, which is expressed as follows:

The parameter used in the SVM can greatly influence recognition performance. To select the optimal parameters, PSO, ACO, GSO, etc., are used; in our study the fruit fly optimization algorithm is used.

5.4 Fruit fly optimization algorithm

The fruit fly optimization algorithm is a technique that is dependent on the fruit fly foraging mannerism for the sake of global optimization (). It is a well-known and intelligent optimization algorithm as it is easy to implement and optimize. The diagram presented in Figure 6 illustrates the FFO algorithm overall.

FIGURE 6

The main steps of the FOA are as follows:

  • Random initialization of the position of the population is conducted. For the population’s position, the P-axis is denoted as an abscissa value, and the Q-axis is denoted as an ordinate value.

  • Random evaluation of the direction and position of flying is conducted for every fruit fly. The new position of every fruit fly is denoted as , where . The number of fruit flies in the population is denoted as .

  • For each fruit fly, its distance of it to the origin of the smell assessment value is computed as follows:

  • To compute the smell concentration value, the smell assessment value is used as a fitness function and is represented as . Across the entire population, the fruit fly with the highest smell assessment is identified and represented as

  • The highest smell assessment value which is the best value along with its position are recorded.

  • By means of implementing vision, the fruit fly population enters this position.

  • The iteration of steps 2 to 4 is conducted continuously in this process.

  • Step 6 can be executed if the smell assessment value is better than the previous one.

6 Results and discussion

A public seismic signal dataset called STEAD, which is available online, was used in our study to conduct the proposed experiment (). This dataset is massive and has seismic signal recordings of approximately 19,000 h, comprising two categories: local earthquake waveform and seismic noise waveforms. A time duration of 60 s is spanned by every signal in the dataset. The dataset has records from 1984 to 2018 and about 450,000 earthquake instances. The resampling of the signals within the dataset was done at 100 Hz, so the data can be harmonized well. The partitioning of the signals was done into specific segments of about 60s each. Three vital components are encompassed by three segments: east, north, and vertical, expressed as X, Y and Z channels. The magnitude of the earthquake ranges from −0.5 to 7.9 in the STEAD dataset. Two vital predefined cases are assigned in the study: i) noise vs. earthquake classification; ii) noise vs. p-wave vs. s-wave classification. The basic extraction of segments with a time duration of 10 s yielded about 15,000 segments identified as noise and about 9,000 segments identified as p-waves and s-waves. The idea of randomly picking only 2000 segments as per was chosen in this study so that the results would be compared with theirs. Therefore, for each class, 2000 segments were chosen randomly so that the scenario comprises three classes: 2000 for noise segments, 2000 for s-wave segments, and 2000 for p-wave segments. Ultimately, two scenarios are prescribed in this dataset and are explained as follows. The first scenario has two classes: noise and event. There are about 2000 signals in the noise category, and about 4,000 signals in the event category, and these seismic signals belong to P and S waves, respectively. The second scenario has three classes—noise, P waves, and S waves—with a total of 6,000 seismic signals and 2000 seismic signals in each category.

Based on the K-best feature choosing model, the optimal number of clusters is determined using the Silhouette method, and the value of K was assigned at 10 so that the clusters can be defined well. For the RL-BN-AENN concept used here, the learning rate is set as 0.01 and the discount factor as 0.5 for the RL aspect of the study. For the AENN concept, the number of estimators is set as 100, the learning rate is set as 0.2, the number of splits of the decision trees is set as 1, and the minimum number of observations per leaf node is set as 0.5 for AdaBoost. For ENN, the learning rate is set as 0.2, the number of epochs is set as 500, and the L2 regularization is used along with the gradient thresholding clipping method. For DE-BC, the population size is set as 100, the mutation factor is set as 0.8, and the crossover rate is set as 0.2 in this experiment. The learning rate is set as 0.1, the number of epochs is set as 500, and the batch size is set as 50 for the DE-BC classifiers. For LDA–FW–RVM, the regularization parameter is set to 0.5, and for the RVM, the kernel used is RBF, and so the kernel width is assigned as 0.2 in the experiment. As far as LDA–FFOA–SVM is concerned, the regularization parameter for SVM is set as 10, and again the kernel used is RBF in this experiment. For FFOA, the swarm size or population size is set as 100, and the maximum number of iterations is set as 300. The lower and upper boundary ranges of the FFOA can range in the values of [0, 1], respectively.

A desktop machine was used with an Intel i5 processor with 5.4 GHz, 32 GB main memory and the Windows 11 Operating System. The programming environment used was MATLAB 2023a. Ten-fold cross-validation was used in this study. The primary performance evaluation metric utilized is classification accuracy:

In Equation 74, the true positive, true negative, false positive, and false negative are represented by TP, TN, FP, and FN, respectively. A careful examination of Table 1 reveals that LLR with K-best feature selection and the AdaBoost classifier gave the highest classification accuracy of 97.43% for scenario 1. Table 2 reveals that RNE with Kriging model feature selection and the LR classifier gave the highest classification accuracy of 87.73% for scenario 2. It is evident from Table 3 that PCCF with MVMD, Kriging model feature selection, and the AdaBoost classifier produces a high classification accuracy of 96.35% for scenario 1. Table 4 shows that PCCF with MVMD and K-best feature selection and AdaBoost classifier produces a high classification accuracy of 89.83% for scenario 2. Table 5 reveals that LLR with Kriging model feature selection and the AdaBoost ENN produces a high classification accuracy of 98.89% for scenario 1, and Table 6 shows that LLR with the Kriging model feature selection technique and AdaBoost-ENN classifier produces a high classification accuracy of 90.46% for scenario 2. It is evident from Table 7 that PCCF with MVMD, Kriging model feature selection, and AdaBoost-ENN classifier produces a high classification accuracy of 97.58% for scenario 1. Table 8 shows that PCCF with MVMD, Kriging model feature selection, and AdaBoost-ENN classifier produces a high classification accuracy of 89.98% for scenario 2, and Table 9 reveals that a high classification accuracy of 94.53% is obtained when GEL and Kriging model feature selection are classified with the DE-BC classifier for scenario 1. From Table 10, it is evident that a high classification accuracy of 88.48% is obtained when GEL and Kriging model feature selection are utilized and classified with the DE-BC classifier for scenario 2. Careful examination of Table 11 reveals that a high classification accuracy of 92.31% is obtained for PCCF with MVMD, K-best feature selection, and the DE-BC classifier for scenario 1. Table 12 reveals that a high classification accuracy of 89.91% is obtained for the CCF with MVMD combination and Kriging model feature selection and the DE-BC classifier for scenario 2. Careful examination of Table 13 shows that a high classification accuracy of 98.87% is attained for RNE with K-best feature selection and the LDA-FFOA-SVM classifier for scenario 1. As shown in Table 14, a high classification accuracy of 92.78% is obtained when RNE, Kriging model feature selection, and the LDA-FFOA-SVM classifier are used for scenario 2. Tables 15, 16 show high classification accuracies of 98.23% and 92.95% for the combination of PCCF and MVMD and Kriging model feature selection and the LDA-FFOA-SVM classifier for scenarios 1 and 2, respectively.

TABLE 1

K-best feature selectionKriging model feature selection
AdaBoostNBCRFDTLRAdaBoostNBCRFDTLR
GEL96.5694.4591.1290.1292.0594.2393.4493.8390.3292.12
LLR97.4393.8992.3491.3493.1195.4794.6892.2291.3393.44
RNE95.5792.0290.6793.9892.3596.6294.7294.4792.6893.45

Performance analysis of GEL, LLR, and RNE with K-best and Kriging model feature selection techniques and machine learning classifiers for scenario 1.

TABLE 2

K-best feature selectionKriging model feature selection
AdaBoostNBCRFDTLRAdaBoostNBCRFDTLR
GEL83.9882.0985.1184.8787.0982.2180.8681.0283.9284.56
LLR85.7781.4583.2380.8385.9885.3483.7884.3484.3585.87
RNE86.5683.7886.6881.4583.9283.5684.9286.7886.7687.73

Performance analysis of GEL, LLR, and RNE with K-best and Kriging model feature selection techniques and machine learning classifiers for scenario 2.

TABLE 3

K-best feature selectionKriging model feature selection
AdaBoostNBCRFDTLRAdaBoostNBCRFDTLR
CCF + MVMD94.3292.1190.1291.2390.2195.3493.4492.4593.4592.45
PCCF + MVMD95.2293.1492.4593.2292.3496.3594.5594.6795.3395.03

Performance analysis of CCF, PCCF, and MVMD with K-best and Kriging model feature selection techniques and machine learning classifiers for scenario 1.

TABLE 4

K-best feature selectionKriging model feature selection
AdaBoostNBCRFDTLRAdaBoostNBCRFDTLR
CCF + MVMD88.7185.0284.2286.4185.9289.0286.6685.0287.6686.02
PCCF + MVMD89.8386.3185.3187.2686.2189.3487.5186.4187.7987.31

Performance analysis of CCF, PCCF, and MVMD with K-best and Kriging model feature selection techniques and machine learning classifiers for scenario 2.

TABLE 5

K-best feature selectionKriging model feature selection
AdaBoostENNAdaBoost–ENNAdaBoostENNAdaBoost–ENN
GEL94.0894.1197.0293.5794.0198.47
LLR93.8894.2598.3394.8895.2198.89
RNE94.9395.7698.4794.9796.3398.88

Performance analysis of GEL, LLR, and RNE with K-best and Kriging model feature selection techniques and AdaBoost, ENN, and AdaBoost–ENN classifiers for Scenario 1.

TABLE 6

K-best feature selectionKriging model feature selection
AdaBoostENNAdaBoost–ENNAdaBoostENNAdaBoost–ENN
GEL84.4586.0987.2284.9886.0289.32
LLR83.5685.8588.3585.2487.4390.46
RNE86.8986.2188.6886.7386.1189.78

Performance analysis of GEL, LLR, and RNE with K-best and Kriging model feature selection techniques and AdaBoost, ENN, and AdaBoost–ENN classifiers for Scenario 2.

TABLE 7

K-best feature selectionKriging model feature selection
AdaBoostENNAdaBoost–ENNAdaBoostENNAdaBoost–ENN
CCF + MVMD92.3493.0395.0293.9494.1296.02
PCCF + MVMD93.3893.4196.0394.8995.5797.58

Performance analysis of CCF, PCCF, and MVMD with K-best and Kriging model feature selection techniques and AdaBoost, ENN, and AdaBoost–ENN classifiers for scenario 1.

TABLE 8

K-best feature selectionKriging model feature selection
AdaBoostENNAdaBoost–ENNAdaBoostENNAdaBoost–ENN
CCF + MVMD85.6786.8889.3186.0887.2189.67
PCCF + MVMD86.7887.0388.9987.5588.4489.98

Performance analysis of CCF, PCCF, and MVMD with K-best and Kriging model feature selection techniques and AdaBoost, ENN, and AdaBoost-ENN classifiers for scenario 2.

TABLE 9

K-best feature selectionKriging model feature selection
Bayesian classifierDE–BCBayesian classifierDE-BC
GEL90.5692.5489.3494.53
LLR91.6694.2190.4693.46
RNE89.0393.4591.3593.78

Performance analysis of GEL, LLR, and RNE with K-best and Kriging model feature selection techniques and Bayesian classifier/DE-BC classifier for scenario 1.

TABLE 10

K-best feature selectionKriging model feature selection
Bayesian classifierDE–BCBayesian classifierDE–BC
GEL83.3486.4585.4588.48
LLR82.1285.1386.2288.22
RNE80.2385.6986.2988.41

Performance analysis of GEL, LLR, and RNE with K-best and Kriging model feature selection techniques and Bayesian classifier/DE–BC classifier for scenario 2.

TABLE 11

K-best feature selectionKriging model feature selection
Bayesian classifierDE–BCBayesian classifierDE–BC
CCF + MVMD89.8191.3490.3492.22
PCCF + MVMD90.3492.3191.2392.01

Performance analysis of CCF, PCCF, and MVMD with K-best and Kriging model feature selection techniques and Bayesian classifier/DE–BC classifier for scenario 1.

TABLE 12

K-best feature selectionKriging model feature selection
Bayesian classifierDE–BCBayesian classifierDE–BC
CCF + MVMD85.3987.8986.8889.91
PCCF + MVMD86.8888.0187.9989.22

Performance analysis of CCF, PCCF, and MVMD with K-best and Kriging model feature selection techniques and Bayesian classifier/DE–BC classifier for scenario 2.

TABLE 13

K-best feature selectionKriging model feature selection
LDA–FW–RVMLDA–FFOA–SVMLDA–FW–RVMLDA–FFOA–SVM
GEL97.6798.2197.0898.22
LLR96.8898.3497.6898.57
RNE96.9098.8796.2297.93

Performance analysis of GEL, LLR, and RNE with K-best and Kriging model feature selection techniques and LDA–FW–RVM and LDA–FFOA–SVM classifiers for scenario 1.

TABLE 14

K-best feature selectionKriging model feature selection
LDA–FW–RVMLDA–FFOA–SVMLDA–FW–RVMLDA–FFOA–SVM
GEL87.5789.2388.2391.08
LLR86.9890.5690.2992.29
RNE88.8591.6191.3692.78

Performance analysis of GEL, LLR, and RNE with K-best and Kriging model feature selection techniques and LDA–FW–RVM and LDA–FFOA–SVM classifiers for scenario 2.

TABLE 15

K-best feature selectionKriging model feature selection
LDA–FW–RVMLDA–FFOA–SVMLDA–FW–RVMLDA–FFOA–SVM
CCF + MVMD95.4396.7796.7797.87
PCCF + MVMD96.5697.8897.8198.23

Performance analysis of CCF, PCCF, and MVMD with K-best and Kriging model feature selection techniques and LDA–FW–RVM and LDA–FFOA–SVM classifiers for scenario 1.

TABLE 16

K-best feature selectionKriging model feature selection
LDA–FW–RVMLDA–FFOA–SVMLDA–FW–RVMLDA–FFOA–SVM
CCF + MVMD89.0290.2191.3692.46
PCCF + MVMD90.0391.3892.1792.95

Performance analysis of CCF, PCCF, and MVMD with K-best and Kriging model feature selection techniques and LDA–FW–RVM and LDA–FFOA–SVM classifiers for scenario 2.

A performance comparison illustration of GEL, LLR, and RNE with feature selection and standard machine learning classifiers for scenario 1 is shown in Figure 7, with the highest classification accuracy reported being 98.89%. A performance comparison illustration of GEL, LLR, and RNE with feature selection and LDA–FW–RVM and LDA–FFOA–SVM classifiers for scenario 1 is shown in Figure 8, with the highest classification accuracy reported being 98.87%. A performance comparison illustration of CCF, PCCF, and MVMD with feature selection and LDA–FW–RVM and LDA–FFOA–SVM classifiers for scenario 1 is shown in Figure 9, with the highest classification accuracy reported being 98.23%. A performance comparison illustration of CCF, PCCF, and MVMD with feature selection and LDA–FW–RVM and LDA–FFOA–SVM classifiers for scenario 2 is shown in Figure 10, with a high classification accuracy reported at 92.95%. A performance comparison illustration of GEL, LLR, and RNE with feature selection and LDA–FW–RVM and LDA–FFOA–SVM classifiers for scenario 2 is shown in Figure 11, with a high classification accuracy reported at 92.78%. A performance comparison illustration of GEL, LLR, and RNE with feature selection and AdaBoost, ENN, and AdaBoost-ENN classifiers for scenario 2 is shown in Figure 12, with a high classification accuracy reported at 90.46%.

FIGURE 7

FIGURE 8

FIGURE 9

FIGURE 10

FIGURE 11

FIGURE 12

6.1 Comparison with previous research

The results from the present study are compared with the results from previously presented work for a comprehensive analysis in Table 17.

TABLE 17

AuthorMethod adoptedClassification typeClassification metric and result analysis (%)
Zhu and Beroza (2019)PhaseNet based CNN3 classes (noise, P waves, S waves)P waves
Precision = 93.9%
S waves
Precision = 85.3%
Saad and Chen (2020)Capsule neural network2 classes (noise, earthquake)Accuracy = 98.61%
Saad and Chen (2021)CapsPhase technique3 classes (noise, P waves, S waves)P waves
Precision = 94.5%
S waves
Precision = 88.05%
Shakeel et al. (2021)Recurrent neural network with 3D CNN2 classes (earthquake, no earthquake)Accuracy = 98.01%
LSTM, CNN with Keras2 classes (noise, earthquake)Accuracy = 99.20%
CNN2 classes (noise, earthquake)Accuracy = 98.80%
Fully convolutional DenseNet2 classes (noise, earthquake)Accuracy = 99.37%
U-shape neural network2 classes (noise, earthquake)Accuracy = 98.80%
Convolutional recurrent model2 classes (noise, earthquake)Accuracy = 98.58%
ANFIS2 classes (earthquake, no earthquake)Precision = 92.52%
Butterfly pattern with machine learning2 classes (noise, earthquake)
3 classes (noise, P waves, S waves)
Accuracy = 99.62%
Accuracy = 93.13%
Proposed researchLLR with Kriging model-based feature selection and AdaBoost-ENN classifier2 classesAccuracy = 98.89%
RNE with K-best feature selection and LDA–FFOA–SVM classifier2 classesAccuracy = 98.87%
PCCF with MVMD and Kriging model feature selection classified with LDA–FFOA–SVM classifier2 classesAccuracy = 98.23%
PCCF with MVMD and Kriging model feature selection classified with LDA–FFOA–SVM classifier3 classesAccuracy = 92.95%
RNE with Kriging model-based feature selection and LDA–FFOA–SVM classifier3 classesAccuracy = 92.78%
LLR with Kriging model-based feature selection and AdaBoost-ENN classifier3 classesAccuracy = 90.46%

Comparison of the current results with previous research.

Most of the models proposed previously have concentrated on a two-class classification model, while this study also concentrates equally on both two- and three-class classification issues. In previous research, only a single model has been proposed and validated, while in this study, four different and efficient strategies are proposed. While the current results show slightly fewer classification results than results previously obtained for the three-class classification issues, it is evident that the models are not very complex overall. The proposed models are quite robust and simple, and they can also be applied to other signal processing models. For the two-class classification problem, the proposed models provided very good results, showing that the proposed methods are quite promising. Computational complexity was assessed as in this experiment. The standard two-sided Wilcoxon analysis is implemented on a statistical test in this experiment, and a good confidence level was obtained for all the features, thus showing its potential and adaptability for classification. The computational time analysis for the proposed studies is shown in Table 18.

TABLE 18

YearAuthorMethodComputational time
2026Proposed researchLLR with Kriging model feature selection and AdaBoost-ENN classifier (2 classes)5.349 s
RNE with K-best feature selection and LDA–FFOA–SVM classifier (2 classes)6. 903 s
PCCF with MVMD and Kriging model feature selection classified with LDA–FFOA–SVM classifier (2 classes)6. 365 s
PCCF with MVMD and Kriging model feature selection classified with LDA–FFOA–SVM classifier (3 classes)7.117 s
RNE with Kriging model feature selection and LDA–FFOA–SVM classifier (3 classes)7.427 s
LLR with Kriging model feature selection and AdaBoost–ENN classifier (3 classes)8.011 s

Computational time analysis for the proposed research.

It is evident from Table 18 that a low computational time of 5.349 s is achieved for the LLR with Kriging model feature selection and AdaBoost–ENN classifier for two classes, and high computational time of 8.011 is achieved for the LLR with Kriging model feature selection and AdaBoost–ENN classifier for three classes.

The major limitation of the proposed models is that there is no specific rationale in such a framework as to why only specific set of algorithms with machine learning concepts are used. Moreover, when machine learning concepts are used, some issues must be carefully managed every time, such as high data dependency, lack of interpretability in some instances, overfitting/underfitting problems, bias issues, lack of generalization, and overdependence on statistical correlations.

7 Conclusion and future research

Previously, the expert knowledge of a seismologist was required for the manual analysis of earthquake data, which was highly time-consuming, and erroneous results have often been reported. However, with the advent of automated algorithms and machine learning techniques, the classification of earthquake seismic signals is now much easier. In this study, four different techniques are proposed for seismic signal classification, with the highest classification accuracy of 98.89% being obtained for the two-class classification problem when LLR with Kriging model feature selection and the AdaBoost–ENN classifier is implemented. A high classification accuracy of 92.95% is obtained for the three-class classification problem when PCCF with MVMD and Kriging model feature selection classified with the LDA–FFOA–SVM classifier is implemented. In this study, only one dataset was analyzed; future research will be implemented on multiple seismic datasets. Moreover, a real time test bed will be built wherein the proposed models can be validated successfully with real time seismic data. A variety of hybrid models will also be incorporated as feature selection techniques, and an exhaustive performance analysis will be conducted to understand the best performing techniques for the classification of seismic signals.

Statements

Data availability statement

Publicly available datasets were analyzed in this study. These data can be found at: the STEAD dataset is available from https://github.com/smousavi05/STEAD, and it can be referred to the paper for further details ‘Mousavi, S. M., Y. Sheng, W. Zhu, and G. C. Beroza (2019), STanford EArthquake Dataset (STEAD): A Global Data Set of Seismic Signals for AI, IEEE Access’.

Author contributions

SP: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Software, Writing – original draft, Writing – review and editing. D-OW: Funding acquisition, Project administration, Resources, Supervision, Validation, Visualization, Writing – original draft, Writing – review and editing.

Funding

The author(s) declared that financial support was received for this work and/or its publication. This research was supported by the Regional Innovation System & Education (RISE) Glocal University 30 Project program through the Gangwon RISE Center, funded by the Ministry of Education (MOE) and the Gangwon State (G.S.), Republic of Korea (2025-RISE-10-009).

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Abbreviations

GEL, graph embedding learning; LRR, low-rank representation; RNE, robust neighborhood embedding; CCF, cross-correlation function; PCCF, partial cross-correlation function; MVMD, multivariate variational mode decomposition; RL, reinforcement learning; BNs, Bayesian networks; ENNs, Elman neural networks; RVM, relevance vector machine; FW, fuzzy weighted; AENN, AdaBoost-based Elman neural network; DE, differential evolution; BC, Bayesian classifier; LDA, linear discriminant analysis; FFOA, fruit fly optimization algorithm; SVM, support vector machine.

References

  • 1

    AbdullahS. M. U.RehmanN. u.KhanM. M.MandicD. P. (2015). A multivariate empirical mode decomposition based approach to pan sharpening. IEEE Trans. Geoscience Remote Sens.53 (7), 39743984. 10.1109/tgrs.2015.2388497

  • 2

    AsankaD.PereraA. S. (2018). “Defining fuzzy membership function for fuzzy data warehouses,” in 2018 4th International Conference for Convergence in Technology (I2CT) (Mangalore, India), 15.

  • 3

    AtmajaB. K.MustikaI. W.HidayatR.AdiH. N.MusaddidA. T. (2022). “Development of CNN pruning method for earthquake signal imagery classification,” in 2022 14th International Conference on Information Technology and Electrical Engineering (ICITEE), Yogyakarta, Indonesia, 18-19 October 2022 (IEEE), 298302.

  • 4

    BatrakovD. O.GolovinD. V.SimachevA. A.BatrakovaA. G. (2010). “Hilbert transform application to the impulse signal processing,” in 2010 5th International Conference on Ultrawideband and Ultrashort Impulse Signals (Sevastopol, Ukraine), 113115.

  • 5

    BejigaM. B.ZeggadaA.NouffidjA.MelganiF. (2017). A convolutional neural network approach for assisting avalanche search and rescue operations with UAV imagery. Remote Sens.9, 100. 10.3390/rs9020100

  • 6

    BhatiaM.AhangerT. A.ManochaA. (2023). Artificial intelligence based real-time earthquake prediction. Eng. Appl. Artif. Intell.120, 105856. 10.1016/j.engappai.2023.105856

  • 7

    BilalM. A.JiY.WangY.AkhterM. P.YaqubM. (2022). Early earthquake detection using batch normalization graph convolutional neural network (BNGCNN). Appl. Sci.12, 7548. 10.3390/app12157548

  • 8

    CaoY. (2010). “Study of the Bayesian networks,” in 2010 International Conference on E-Health Networking Digital Ecosystems and Technologies (EDT) (Shenzhen, China), 172174.

  • 9

    ChakrabortyM.FennerD.LiW.FaberJ.ZhouK.RümpkerG.et al (2022). CREIME—A convolutional recurrent model for earthquake identification and magnitude estimation. J. Geophys. Res. Solid Earth127, e2022JB024595.

  • 10

    DemirA. (2025). Post-earthquake structural damage assessment, lessons learned, and addressing objections following the 2023 Kahramanmaras, Turkey earthquakes. Bull. Earthq. Eng.23, 11071127. 10.1007/s10518-024-02092-8

  • 11

    EguchiR. T.HuyckC. K.GhoshS.AdamsB. J.McMillanA. (2010). “Utilizing new technologies in managing hazards and disasters,” in Geospatial Techniques in Urban Hazard and Disaster Analysis. Editors ShowalterP. S.LuY. (Dordrecht, Netherlands: Springer), 295323.

  • 12

    ElsayedH. S.SaadO. M.SolimanM. S.ChenY.YounessH. A. (2022). Attention-based fully convolutional DenseNet for earthquake. Detect. IEEE Trans. Geoscience Remote Sens.60, 110. 10.1109/tgrs.2022.3194196

  • 13

    FurutaniT.MinamiM. (2021). “Drones for disaster risk reduction and crisis response,” in Emerging Technologies for Disaster Resilience: Practical Cases and Theories. Editors SakuraiM.ShawR. (Singapore: Springer), 5162.

  • 14

    GaoX.ZhangK.TaoD.LiX. (2012). Image super-resolution with sparse neighbor embedding. IEEE Trans. Image Process.21 (7), 31943205. 10.1109/TIP.2012.2190080

  • 15

    GoebelT. H.BrodskyE. E. (2018). The spatial footprint of injection Wells in a global compilation of induced earthquake sequences. Science361, 899904. 10.1126/science.aat5449

  • 16

    IjazH.AhmadR.AhmedR.AhmedW.KaiY.JunW. (2023). A UAV-assisted edge framework for real-time disaster management. IEEE Trans. Geosci. Remote Sens.61, 1001013. 10.1109/tgrs.2023.3306151

  • 17

    IshakL. M. (2017). “Partial discharge location algorithm based on cross-correlation technique for unsynchronized measurement,” in 2017 IEEE 15th Student Conference on Research and Development (Scored) (Wilayah Persekutuan Putrajaya, Malaysia), 388391.

  • 18

    JiaoJ.VenkatK.HanY.WeissmanT. (2015). “Maximum likelihood estimation of information measures,” in 2015 IEEE International Symposium on Information Theory (ISIT) (Hong Kong, China), 839843.

  • 19

    JozinovićD.LomaxA.ŠtajduharI.MicheliniA. (2020). Rapid prediction of earthquake ground shaking intensity using raw waveform data and a convolutional neural network. Geophys. J. Int.222, 13791389. 10.1093/gji/ggaa233

  • 20

    KaurK.SalgotraR.SinghU. (2017). “An improved firefly algorithm for numerical optimization,” in 2017 International Conference on Innovations in Information, Embedded and Communication Systems (ICIIECS) (Coimbatore, India), 15.

  • 21

    KerleN.NexF.DuarteD.VetrivelA. (2019). Uav-based structural damage mapping—results from 6 years of research in two European projects. Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci., 187194. 10.5194/isprs-archives-xlii-3-w8-187-2019

  • 22

    KhanI.KwonY.-W. (2022). P-detector: real-time P-wave detection in a seismic waveform recorded on a low-cost MEMS accelerometer using deep learning. IEEE Geoscience Remote Sens. Lett.19, 15. 10.1109/lgrs.2022.3161017

  • 23

    KhanI.ChoiS.KwonY.-W. (2020). Earthquake detection in a static and dynamic environment using supervised machine learning and a novel feature extraction method. Sensors20, 800. 10.3390/s20030800

  • 24

    KhanI.PandeyM.KwonY.-W. (2021). “An earthquake alert system based on a collaborative approach using smart devices,” in 2021 IEEE/ACM 8th International Conference on Mobile Software Engineering and Systems (Mobilesoft) (IEEE), 6164.

  • 25

    KimK. (2019). Situating the anthropocene: the social construction of the pohang ‘triggered’ earthquake. J. Sci. Technol. Stud.19, 51117.

  • 26

    LeónS. M.CalviñoB. O.VivasL. A.CorretgerR. C.UlacioO. R. (2022). Small-layered feed-forward and convolutional neural networks for efficient P wave earthquake. Detection Expert Syst. Appl.206.

  • 27

    LiY.GuoS.ZhuL.MukaiT.GanZ. (2019). Enhanced probabilistic inference algorithm using probabilistic neural networks for learning control. IEEE Access7, 184457184467. 10.1109/access.2019.2959876

  • 28

    LiW.ChakrabortyM.FennerD.FaberJ.ZhouK.RümpkerG.et al (2022). EPick: attention-based multi-scale UNet for earthquake detection and seismic phase picking. Front. Earth Sci.10, 2075. 10.3389/feart.2022.953007

  • 29

    LiJ.LiK.TangS. (2023). Automatic arrival-time picking of P-and S-waves of microseismic events based on object detection and CNN. Soil Dyn. Earthq. Eng.164, 107560. 10.1016/j.soildyn.2022.107560

  • 30

    LiuH.ZhangZ.ZhangT. (2022). Shale gas investment decision–making: green and efficient development under market, technology and environment uncertainties. Appl. Energy306, 118002. 10.1016/j.apenergy.2021.118002

  • 31

    MajstorovićJ.Giffard-RoisinS.PoliP. (2021). Designing convolutional neural network pipeline for near-fault earthquake catalog extension using single-station waveforms. J. Geophys. Res. Solid Earth126, e2020JB021566. 10.1029/2020jb021566

  • 32

    MaoY.HaoY.CaoX.FangY.LinX.MaoH.et al (2024). Dynamic graph embedding via meta-learning. IEEE Trans. Knowl. Data Eng.36 (7), 29672979. 10.1109/tkde.2023.3329238

  • 33

    MarkopoulosP. P. (2017). “Linear discriminant analysis with few training data,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (New Orleans, LA, USA), 46264630.

  • 34

    MengB. (2018). “Improved elman neural network and its application,” in 2018 Chinese Control and Decision Conference (CCDC) (Shenyang, China), 320324.

  • 35

    MousaviS. M.ShengY.ZhuW.BerozaG. C. (2019). Stanford earthquake dataset (STEAD): a global data set of seismic signals for AI. IEEE Access.

  • 36

    MurtiM. A.JuniorR.AhmedA. N.ElshafieA. (2022). Earthquake multi-classification detection-based velocity and displacement data filtering using machine learning algorithms. Sci. Rep.12, 21200. 10.1038/s41598-022-25098-1

  • 37

    MustikaI. W.AdiH. N.NajibF. (2021). “Comparison of keras optimizers for earthquake signal classification based on deep neural networks,” in 2021 4th International Conference on Information and Communications Technology (ICOIACT) (IEEE), 304308.

  • 38

    OzkayaS. G.BayginM.BaruaP. D. (2024). An automated earthquake classification modal based on a new butterfly pattern using seismic signals. Expert Syst. Appl.238 (Part D).

  • 39

    PriceK. V. (1996). “Differential evolution: a fast and simple numerical optimizer,” in Proceedings of North American Fuzzy Information Processing (Berkeley, CA, USA), 524527.

  • 40

    QiangW.ZhongliZ. (2011). “Reinforcement learning model, algorithms and its application,” in 2011 International Conference on Mechatronic Science, Electric Engineering and Computer (MEC) (Jilin, China), 11431146.

  • 41

    SaadO. M.ChenY. (2020). Earthquake detection and P-wave arrival time picking using capsule neural network. IEEE Trans. Geoscience Remote Sens.59, 62346243. 10.1109/tgrs.2020.3019520

  • 42

    SaadO. M.ChenY. (2021). CapsPhase: capsule neural network for seismic phase classification and picking. IEEE Trans. Geoscience Remote Sens.60, 111. 10.1109/tgrs.2021.3089929

  • 43

    SchweierC.MarkusM. (2006). Classification of collapsed buildings for fast damage and loss assessment. Bull. Earthq. Eng.4, 177192. 10.1007/s10518-006-9005-2

  • 44

    SefatM. S.ShahjahanM.RahmanM.VallesD. (2021). “Ensemble training with classifiers selection mechanism,” in 2021 IEEE 12th Annual Ubiquitous Computing, Electronics & Mobile Communication Conference (UEMCON) (New York, NY, USA), 01310136.

  • 45

    ShakeelM.ItoyamaK.NishidaK.NakadaiK. (2021). Detecting earthquakes: a novel deep learning-based approach for effective disaster response. Appl. Intell.51, 83058315. 10.1007/s10489-021-02285-7

  • 46

    TaunkK.DeS.VermaS.SwetapadmaA. (2019). “A brief review of nearest neighbor algorithm for learning and classification,” in 2019 International Conference on Intelligent Computing and Control Systems (ICCS) (Madurai, India), 12551260.

  • 47

    TehseenR.FarooqM. S.AbidA. (2021). A framework for the prediction of earthquake using federated learning. PeerJ Comput. Sci.7, e540. 10.7717/peerj-cs.540

  • 48

    WangY.YinH.ZhangT.YinM. (2022). “Application research of kriging model in modeling and simulation of aircraft industry process,” in MEMAT 2022; 2nd International Conference on Mechanical Engineering, Intelligent Manufacturing and Automation Technology (Guilin, China), 15.

  • 49

    WangQ. (2022). “Support vector machine algorithm in machine learning,” in 2022 IEEE International Conference on Artificial Intelligence and Computer Applications (ICAICA) (Dalian, China), 750756.

  • 50

    WangQ.YeX.WangN. (2024). Learning low-rank representation approximation for few-shot deep subspace clustering. IEEE Trans. Circuits Syst. Video Technol.34 (11), 1059010603. 10.1109/tcsvt.2024.3411615

  • 51

    WangH.YangS.FludeS.ZhangK.ZhangQ.ZhaoJ.et al (2026). Development, challenges and potentials of underground energy storage systems. Tunn. Undergr. Space Technology170, 107360. 10.1016/j.tust.2025.107360

  • 52

    WuA.LeeJ.KhanI.KwonY.-W. (2021). “CrowdQuake+: data-driven earthquake early warning via IoT and deep learning,” in 2021 IEEE International Conference on Big Data (Big Data), Orlando, FL, USA, 15-18 December 2021 (IEEE), 20682075.

  • 53

    XuG.ShanP.LaiX.HuQ.YangS.XuH. (2026a). Thermomechanical behaviour and damage mechanism of the lining backfill body of high-temperature thermal energy storage reservoirs in mines. Tunn. Undergr. Space Technology171, 107409. 10.1016/j.tust.2025.107409

  • 54

    XuG.LiK.ZhangB.ZhanR.YangS.XiX.et al (2026b). Thermo-mechanical damage characteristics and fracture behavior of backfill bodies in thermal energy storage cavities: an experimental and numerical study. Tunn. Undergr. Space Technology175, 107715. 10.1016/j.tust.2026.107715

  • 55

    YangX.LiuW.LiuW.TaoD. (2021). A survey on canonical correlation analysis. IEEE Trans. Knowl. Data Eng.33 (6), 23492368. 10.1109/tkde.2019.2958342

  • 56

    ZerroukiN.HarrouF.SunY.HouacineA. (2017). “Adaboost-based algorithm for human action recognition,” in 2017 IEEE 15th International Conference on Industrial Informatics (INDIN) (Emden, Germany), 189193.

  • 57

    ZhangX.FanM.WangD.ZhouP.TaoD. (2021). Top-k feature selection framework using robust 0–1 integer programming. IEEE Trans. Neural Netw. Learn. Syst.32 (7), 30053019. 10.1109/tnnls.2020.3009209

  • 58

    ZhuW.BerozaG. C. (2019). PhaseNet: a deep-neural-network-based seismic arrival-time picking method. Geophys. J. Int.216, 261273.

  • 59

    ZhuW.TaiK. S.MousaviS. M.BailisP.BerozaG. C. (2022). An end-to-end earthquake detection method for joint phase picking and association using deep learning. J. Geophys. Res. Solid Earth127, e2021JB023283. 10.1029/2021jb023283

Summary

Keywords

classification, feature extraction, feature selection, optimization, reinforcement learning, support vector machine

Citation

Prabhakar SK and Won D-O (2026) A framework for earthquake seismic signal classification with the aid of hybrid feature selection schemes and robust machine learning models. Front. Earth Sci. 14:1748893. doi: 10.3389/feart.2026.1748893

Received

27 November 2025

Revised

06 May 2026

Accepted

13 May 2026

Published

10 July 2026

Volume

14 - 2026

Edited by

Sheng Nie, Chinese Academy of Sciences (CAS), China

Reviewed by

Pengfei Shan, Xi’an University of Science and Technology, China

Shimaa Elkhouly, National Research Institute of Astronomy and Geophysics, Egypt

Updates

Copyright

*Correspondence: Dong-Ok Won,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics