ORIGINAL RESEARCH article

Front. Earth Sci., 30 July 2026

Sec. Geoinformatics

Volume 14 - 2026 | https://doi.org/10.3389/feart.2026.1822287

A distribution-aware Riemannian transformer for intelligent simulation and early warning of risks in geological energy and underground spaces

  • School of Big Data, Shanxi Finance and Taxation College, Taiyuan, China

Abstract

Intelligent simulation and early warning of risks in geological energy development and underground space utilization are essential for improving resource assessment accuracy, operational safety, and sustainable subsurface engineering. However, geological energy and underground space data, such as seismic images, monitoring maps, simulation fields, and engineering inspection images, usually exhibit heterogeneous distributions, complex spatial correlations, strong noise interference, and non-Euclidean structural dependencies. Conventional Euclidean deep learning models often fail to capture these intrinsic geometric and statistical characteristics, limiting their reliability in risk identification and early-warning tasks. To address these challenges, this paper proposes RGSTNet, a Distribution-Aware Riemannian Transformer for intelligent simulation-oriented classification and risk early warning in geological energy and underground spaces. RGSTNet integrates three complementary components: a Riemannian Geometry Module for manifold-based structural representation learning, a Statistical Feature Enhancement Module for big-data-driven distribution-aware reweighting and normalization, and a Manifold-Aware Transformer for global feature fusion under non-Euclidean metrics. Moreover, a joint optimization objective combining label smoothing, Riemannian contrastive learning, and MMD-Fréchet distribution alignment is designed to improve classification accuracy, geometric consistency, and statistical robustness. Experimental results on both general image benchmarks and the Netherlands F3 seismic facies dataset demonstrate that RGSTNet outperforms representative CNN-, graph-, and Transformer-based baselines in Precision, Recall, Accuracy, and F1-score, achieving substantial improvements on domain-relevant geological data. These results indicate that RGSTNet provides a promising distribution-aware deep learning framework for intelligent simulation interpretation, risk classification, and early warning in geological energy resources and underground spaces.

1 Introduction

The intelligent simulation and early warning of risks in geological energy resources and underground spaces have become critical research topics under the background of global energy transition, carbon neutrality, and rapid urbanization. Geological energy resources, including oil, natural gas, geothermal energy, underground gas storage, hydrogen storage, and compressed-air energy storage, are increasingly being developed in deeper and more complex geological environments. Meanwhile, the large-scale construction and utilization of underground spaces, such as urban underground utility tunnels, deep storage caverns, underground transportation systems, and subsurface engineering facilities, require higher standards for structural stability, environmental perception, and dynamic risk prevention. In these scenarios, massive geological, geophysical, monitoring, and simulation data are continuously generated from seismic exploration, well logging, remote sensing, geological mapping, numerical simulation, and Internet-of-Things sensing systems. These large-scale data provide an important foundation for resource evaluation, subsurface structure interpretation, engineering state identification, intelligent simulation analysis, and risk early warning. However, geological energy and underground space data often exhibit complex geometric structures, heterogeneous statistical distributions, multi-scale spatial dependencies, and non-Euclidean feature relationships, which are difficult to capture using conventional Euclidean deep learning models (). For example, seismic facies, reservoir heterogeneity, fracture networks, geothermal anomaly fields, tunnel deformation images, and underground monitoring maps usually contain irregular spatial correlations and latent manifold structures. Traditional feature extraction methods and standard neural networks may fail to fully describe these intrinsic geological and engineering relationships, thereby limiting the accuracy and robustness of simulation interpretation, risk classification, and early-warning decision support (). Therefore, developing a big-data-driven and statistically robust intelligent framework that can jointly model geometric structures and feature distributions has become a key challenge for geological energy development and underground space safety management (; ).

In recent years, deep learning and artificial intelligence have been increasingly applied to image classification, geological interpretation, engineering monitoring, numerical simulation analysis, and safety assessment tasks (; ). On the one hand, convolutional neural network (CNN)-based models can efficiently extract local spatial features from geological images, seismic sections, monitoring maps, and simulation results through local receptive fields and weight-sharing mechanisms. These models have provided useful tools for lithofacies recognition, reservoir classification, fracture identification, tunnel defect detection, deformation monitoring, and underground engineering condition assessment. However, CNNs are generally limited in modeling long-range spatial dependencies and global structural correlations. This limitation becomes more significant when geological or engineering data involve complex stratigraphic structures, discontinuous faults, heterogeneous reservoir properties, or distributed deformation patterns in underground spaces (). By contrast, Transformer models can capture global dependencies through self-attention mechanisms and have shown strong potential in visual representation learning, large-scale data modeling, and simulation result analysis (; ). Nevertheless, conventional Transformers are usually formulated under Euclidean assumptions and rely on Euclidean distance or inner-product similarity to measure feature relationships. Such assumptions may be insufficient for subsurface data with non-Euclidean geometric characteristics, such as curved geological interfaces, irregular fracture networks, manifold-like seismic attributes, and nonlinear coupling relationships among geological, geophysical, hydrological, and engineering variables (). In addition, statistical learning methods have been introduced to improve feature reliability by modeling mean, variance, covariance, and probability density distributions. However, many existing statistical enhancement strategies are separated from geometric modeling and therefore cannot fully exploit the complementary value of geometric structure and statistical regularity in geological big data (). Geometric deep learning offers a promising direction for addressing this limitation by extending deep learning from Euclidean space to graphs, manifolds, and other non-Euclidean domains, thereby providing new tools for complex geological and underground-space data representation ().

Geometric deep learning, as an interdisciplinary field combining deep learning, graph theory, manifold learning, and differential geometry, provides a strong theoretical foundation for modeling irregular and non-Euclidean data structures (). In the field of geological energy resources and underground spaces, many data objects naturally contain geometric and topological relationships. For instance, seismic attribute maps reflect stratigraphic continuity and discontinuity, well-logging curves encode vertical geological variations, reservoir parameter fields exhibit spatial covariance structures, and underground engineering monitoring data reveal deformation, seepage, stress redistribution, and multi-field coupling patterns. These properties suggest that subsurface data should not always be treated as independent Euclidean grids, but should be represented in a geometry-aware and distribution-aware manner. Riemannian geometry is particularly suitable for this purpose because it can describe curved spaces, geodesic distances, tangent-space mappings, and manifold-valued statistical structures (). In practical geological and engineering applications, symmetric positive definite (SPD) matrices can be used to represent covariance, correlation, and local structural characteristics of multi-source data, while geodesic distances can measure intrinsic differences between geological or engineering states more faithfully than Euclidean distances (). Meanwhile, statistical analysis plays an equally important role in geological big data because resource occurrence, reservoir properties, geohazard indicators, simulation outputs, and underground-space monitoring signals often show uncertainty, distribution shift, abnormal fluctuation, and noise disturbance. Statistical quantities such as Fréchet mean, covariance matrices, probability density functions, and distribution alignment metrics can be used to evaluate feature confidence, suppress abnormal observations, and improve model robustness (). Therefore, the integration of Riemannian geometric modeling, big-data-driven statistical enhancement, and deep neural representation learning has the potential to provide a more reliable intelligent framework for simulation interpretation, risk classification, and early warning in complex subsurface environments (; ).

Although existing deep learning methods have achieved promising performance in image classification and data-driven geological interpretation, several key problems remain unresolved when they are applied to intelligent simulation and early warning of risks in geological energy resources and underground spaces. First, most existing models are still based on Euclidean feature spaces, making it difficult to accurately represent non-Euclidean structures such as curved stratigraphic surfaces, irregular fractures, reservoir heterogeneity, tunnel deformation fields, and coupled geological-engineering relationships. Second, many models focus primarily on feature extraction while neglecting the statistical distribution characteristics of large-scale geological, monitoring, and simulation data, resulting in insufficient robustness under noise, outliers, and distribution shifts. Third, conventional attention mechanisms usually calculate feature similarity using Euclidean metrics, which may distort manifold-structured geological features and weaken global feature fusion. Fourth, geological energy development and underground-space utilization involve high safety requirements, while existing methods often lack a joint optimization mechanism that can simultaneously constrain classification accuracy, geometric consistency, and statistical distribution rationality. These limitations restrict the practical application of intelligent models in high-precision resource assessment, underground engineering state identification, simulation-driven risk analysis, and early-warning decision support.

To address these issues, this paper proposes a Distribution-Aware Riemannian Transformer, termed RGSTNet, for intelligent simulation-oriented classification and early warning of risks in geological energy and underground spaces. The proposed framework integrates big-data-driven statistical modeling, manifold-based geometric representation learning, and global feature fusion within an end-to-end architecture. Specifically, the Riemannian Geometry Module first converts input data, such as geological images, seismic attribute maps, underground monitoring images, or simulation-derived feature maps, into structured graph representations. It then embeds local features into Riemannian manifold spaces, especially SPD manifolds, and performs geometric analysis and Riemannian convolution to extract intrinsic non-Euclidean structural information. The Statistical Feature Enhancement Module further models the distribution characteristics of manifold features by using Fréchet mean, covariance estimation, probability density functions, adaptive reweighting, and Riemannian normalization. This module strengthens high-confidence geological and engineering features while suppressing noisy or abnormal patterns, thereby improving the statistical robustness of the learned representation. The Manifold-Aware Transformer then conducts global feature fusion using attention mechanisms adapted to Riemannian metrics, enabling the model to capture long-range dependencies among geological structures, reservoir characteristics, underground-space states, simulation patterns, and safety-related indicators. In addition, a joint loss function combining label smoothing loss, Riemannian contrastive loss, and MMD+Fréchet distribution alignment is designed to jointly optimize classification accuracy, manifold structure consistency, and statistical distribution stability. Through this design, RGSTNet provides a general intelligent framework for geological energy resource evaluation, subsurface feature classification, underground space condition recognition, simulation result interpretation, risk classification, and early warning.

The main contributions of this paper are as follows:

  • We propose RGSTNet, a Distribution-Aware Riemannian Transformer for intelligent simulation-oriented classification and early warning of risks in geological energy resources and underground spaces, integrating Riemannian geometric modeling, big-data-driven statistical feature enhancement, and manifold-aware global attention.

  • RGSTNet consists of three core modules for non-Euclidean structural representation learning, statistical distribution-aware feature optimization, and global manifold feature fusion, enabling robust modeling of heterogeneous geological, engineering, monitoring, and simulation data.

  • We introduce a joint loss function combining label smoothing, Riemannian contrastive loss, and MMD+Fréchet distribution alignment to improve classification accuracy, geometric consistency, statistical robustness, and generalization ability under distribution shifts.

  • We design Riemannian convolution, attention, and normalization operations to support effective learning in manifold space, providing a methodological foundation for intelligent geological resource evaluation, underground space safety assessment, simulation interpretation, risk identification, and early warning.

2 Related works

2.1 Deep learning for image classification

In recent years, deep learning for image classification has witnessed the emergence of a series of representative models and architectures that have substantially improved both classification performance and generalization ability. Traditional convolutional neural network (CNN)-based models, despite many subsequent refinements, still serve as important baselines. For example, EfficientNetV2 () achieves higher parameter efficiency and stronger feature extraction capability, making it well suited to fine-tuning on small- and medium-scale datasets. ConvNeXt (), in contrast, revisits CNN design under modern training strategies and has demonstrated performance comparable to, or even better than, Transformer-based models on multiple datasets.

At the same time, Vision Transformer (ViT) ()-based models have continued to evolve rapidly. By partitioning images into tokens and leveraging self-attention to capture global dependencies, these models perform particularly well in large-scale pretraining and high-resolution image classification. Improved variants, such as the Hierarchical Multi-Scale Vision Transformer, have shown strong performance in medical image classification by enhancing the modeling of features across different spatial scales (). In addition, architectures designed for fine-grained classification have also attracted considerable attention. For instance, the Gradient Focal Transformer (GFT) () uses gradient-based attention to adaptively focus on more discriminative local regions, thereby effectively improving recognition accuracy in complex scenarios.

Beyond CNNs and Transformers, several other directions have also achieved notable progress. ClusterViG (), which represents the development of Vision Graph Neural Networks (Vision GNNs), combines dynamic and efficient graph convolutions with image partition strategies in a hybrid CNN/GNN architecture, yielding clear advantages in both global context modeling and inference efficiency. For resource-constrained environments, RapidNet () introduces a lightweight backbone based on multi-layer dilated convolutions and achieves a favorable trade-off between accuracy and latency on mobile devices. In addition, Vision Mamba/Vim models () represent an attempt to introduce state-space models into visual backbones. By employing bidirectional state-space modules, these methods reduce computational and memory overhead while maintaining competitive classification performance, especially for long-sequence and high-resolution visual representation learning.

Another important trend is the development of visual-language models that integrate visual and linguistic information. For example, pre-trained visual representations derived from DINOv2/CLIP (; ), when combined with Transformer architectures, have demonstrated strong cross-modal generalization in zero-shot and few-shot classification tasks.

However, despite the impressive progress achieved by these emerging models in image classification, several challenges remain. First, although Vision Transformers and their variants, such as the Hierarchical Multi-Scale Vision Transformer, have shown excellent performance on large-scale datasets, their high computational cost and memory consumption remain major barriers to practical deployment (). This issue is particularly critical in resource-constrained settings, where reducing computational burden while preserving high accuracy is still difficult. Second, although Graph Neural Networks (GNNs) have shown promise in capturing complex structural relationships within images, their dependence on graph construction may increase the risk of overfitting when applied to conventional image data, and their scalability to large datasets still requires further improvement ().

Moreover, although lightweight architectures such as RapidNet perform well on mobile devices, their performance on large-scale datasets is often less competitive, which limits their broader applicability (). Similarly, although visual-language models exhibit cross-modal generalization ability in zero-shot classification tasks, effectively handling heterogeneity across modalities, especially under noisy conditions or complex backgrounds, remains an open problem. Therefore, an important current research direction is to explore how to achieve better computational efficiency, stronger generalization, and more effective multimodal fusion in model design so as to address increasingly complex real-world applications (; ).

2.2 Geometric deep learning in image classification

Recent advances in geometric deep learning (GDL) () have significantly expanded the methodological landscape of image classification, particularly for modeling complex non-Euclidean relationships in visual data. One representative direction is manifold learning, which improves feature representation by exploring the low-dimensional manifold structure underlying image data and thereby better captures its geometric characteristics. Geodesic CNNs extend conventional CNNs by incorporating geodesic distances, enabling more effective processing of data defined on curved spaces such as Riemannian manifolds. In addition, Graph Neural Networks (GNNs) have been widely introduced into image classification, with methods such as ClusterViG integrating CNNs and GNNs to model both local and global relationships in images ().

Other geometric approaches have also received increasing attention. Spatial Graph Convolutional Networks (SGCN) (), for example, represent pixel-level relationships as graph structures, thereby facilitating a more explicit characterization of spatial dependencies. Hyperbolic Neural Networks exploit hyperbolic geometry to model hierarchical relationships in data and are therefore particularly suitable for image classification tasks involving complex structures (). In addition, Riemannian Attention Networks (RANs) extend attention mechanisms to non-Euclidean spaces, improving the ability to capture long-range dependencies in geometrically structured data. Deep Set Networks, meanwhile, employ permutation-invariant functions to aggregate feature representations over sets and provide a robust solution for classifying unordered data such as point clouds ().

Building on the above advances in geometric deep learning, this paper further addresses the fragmented treatment of geometric structure modeling and statistical feature fusion in existing methods. We propose an end-to-end framework, RGSTNet (), which integrates Riemannian geometric modeling, statistical feature optimization, and a manifold-aware Transformer. Unlike Geodesic CNNs, which mainly emphasize local geometric feature extraction, RGSTNet converts images into structured representations through the Riemannian Geometry Module and embeds them into manifold spaces such as SPD manifolds. By combining core geometric quantities such as geodesic distance and sectional curvature, the proposed framework explores the intrinsic geometric structure of the data more thoroughly.

Compared with graph-based modeling strategies such as those adopted in ClusterViG and SGCN, our method employs an adaptive k-nearest-neighbor strategy to construct graph connections and performs local feature aggregation in manifold space through Riemannian convolution. This design helps avoid the structural distortion that may arise when non-Euclidean data are processed using Euclidean convolutions. In addition, whereas hyperbolic neural networks mainly focus on hierarchical relationship representation, RGSTNet further introduces a Statistical Feature Enhancement Module. This module performs adaptive reweighting and normalization based on the Fréchet mean, covariance matrices, and probability density functions (PDFs) on the manifold, thereby enabling a deeper integration of geometric structure and statistical information.

Furthermore, compared with the attention mechanism design in Riemannian Attention Networks (RANs), the manifold-aware Transformer proposed in this paper computes attention scores using query-key-value projections adapted to Riemannian space together with geodesic distance metrics. This design alleviates the feature distortion caused by the dependence of conventional Transformers on Euclidean distance, while strengthening global feature fusion through multi-head attention and Riemannian feedforward networks. Finally, unlike the permutation-invariant aggregation mechanism of Deep Set Networks, RGSTNet introduces a joint loss function that combines label smoothing loss, Riemannian contrastive loss, and MMD+Fréchet loss. This formulation not only improves classification accuracy but also constrains geometric structure consistency and statistical distribution rationality, ultimately enhancing both generalization ability and robustness in complex image classification tasks.

3 Methods

3.1 Proposed network

To support intelligent simulation interpretation and risk early warning in geological energy resources and underground spaces, this paper proposes RGSTNet (Figure 1), a Distribution-Aware Riemannian Transformer framework designed to model complex structural and statistical characteristics of image-like subsurface data. In practical geological and underground engineering scenarios, data such as seismic sections, geological images, tunnel inspection images, monitoring maps, and numerical simulation fields often contain irregular spatial dependencies, heterogeneous distributions, noise interference, and non-Euclidean feature relationships. Unlike conventional visual classification models that operate purely in Euclidean space, RGSTNet explicitly considers manifold-valued structures and distribution shifts that may exist in geological and underground-space data. The framework consists of three key components: the Riemannian Geometry Module, the Statistical Feature Enhancement Module, and the Manifold-Aware Transformer.

FIGURE 1

The overall workflow is as follows. First, the input image or image-like simulation field is transformed into a structured graph representation, where local visual regions, pixels, or patches are regarded as nodes and their relationships are modeled through adaptive neighborhood construction. Second, the graph-based representations are embedded into Riemannian manifold spaces, allowing the model to capture intrinsic geometric dependencies in geological structures or underground engineering patterns. Third, statistical feature enhancement is performed to strengthen high-confidence discriminative information and suppress noisy or abnormal feature patterns. Finally, the Manifold-Aware Transformer conducts global feature fusion and produces the final classification or risk-state prediction result. This design makes RGSTNet suitable for general image classification as well as domain-oriented tasks such as seismic facies identification, tunnel defect recognition, simulation field classification, and underground risk early warning.

3.2 Riemannian Geometry Module

Riemannian Geometry Module, as the core foundational component of RGSTNet, has the primary mission of transforming the raw input image into structured data on a Riemannian manifold. Through the progressive steps of graph structure construction, SPD manifold modeling, in-depth geometric feature analysis, and Riemannian convolution operations, it fully exploits the intrinsic geometric structure and relational information of the data, providing strong geometric feature support for the subsequent feature optimization and fusion modules. The operation logic of the entire module is tightly connected, forming a complete feedback loop from data structure transformation to geometric information extraction.

3.2.1 Graph construction

The mathematical formulation of the proposed RGSTNet framework is presented sequentially in Equations 136. To capture the spatial relationships and local topological structure between image pixels, the image must first be converted into structured graph data. Let the input image be , where is the height, is the width, and is the number of channels. Each pixel position corresponds to a feature vector , which is defined as a graph node. The set of nodes is , and the total number of nodes is . To construct the connectivity between nodes, an adaptive -nearest neighbor strategy is used. For each node (where is the node index, ), the cosine distance and Euclidean distance are combined to form a weighted distance with other nodes :where is the distance weighting factor, is a regularization term to avoid division by zero, and denotes the norm. Based on this distance, the -nearest neighbors for each node are selected to construct the edge set , and the graph adjacency matrix is defined, with the elements quantifying the connection strength through a Gaussian kernel function:where is the bandwidth of the Gaussian kernel, and represents the set of -nearest neighbors for node . Through this process, the image is transformed into a graph data structure with a clear topological structure and quantized connection weights, providing structured support for subsequent manifold modeling.

3.2.2 SPD manifold

Symmetric Positive Definite (SPD) manifolds are core spaces for geometric modeling as they can accurately describe the covariance structure and geometric distribution properties of data. The process of mapping graph node features to the SPD manifold consists of two steps: local covariance computation and positive-definite correction. For a node and its neighborhood (including itself), the weighted local covariance matrix is first computed:where is the weighted neighborhood mean, is the neighborhood size, and is a regularization term to avoid matrix singularity. Since the covariance matrix may not be positive-definite due to computational errors, an eigenvalue correction is applied to ensure its SPD property. The corrected SPD matrix is:where is the minimum eigenvalue threshold, is a global regularization coefficient, and is the identity matrix. The corrected matrix satisfies and for any non-zero vector , i.e., . All nodes correspond to points in the SPD manifold, with the set forming the embedded SPD manifold.

3.2.3 Geometric analysis

To quantify the intrinsic geometric properties of features on the SPD manifold, three core geometric quantities need to be calculated: geodesic distance, sectional curvature, and Riemannian gradient. The geodesic distance measures the shortest path between two points on the manifold. It is defined based on the Riemannian logarithm and exponential mappings:where denotes the matrix logarithm (achieved through eigenvalue decomposition: ), and denotes the Frobenius norm. and are the square root and inverse square root matrices of , respectively. The sectional curvature reflects the curvature of the manifold along a specific two-dimensional section. For and orthogonal unit vectors in its tangent space , the sectional curvature is computed as:where denotes the trace of a matrix. The Riemannian gradient provides the direction for function optimization on the manifold. For a target function , the Riemannian gradient is the projection of the Euclidean gradient onto the tangent space:where is the Euclidean gradient of at , and this equation ensures the gradient vector belongs to the tangent space by symmetrizing it.

3.2.4 Riemannian convolution

Traditional convolution relies on the Euclidean space translation invariance and cannot be directly applied to SPD manifolds. Therefore, a Riemannian convolution is designed to perform local feature aggregation on the manifold. This process consists of three steps: tangent space mapping, weighted convolution, and manifold mapping back. First, the neighbor nodes are projected into the tangent space of the center node using the logarithmic mapping:

Then, a weighted convolution operation is performed in the tangent space, with a learnable convolution kernel and bias (where is a symmetric matrix), and the convolution result is:

This equation performs feature transformation through matrix multiplication, using the adjacency weight to quantify the contribution of neighbor features. Finally, the convolution result in the tangent space is mapped back to the SPD manifold using the exponential map, resulting in the Riemannian convolution feature:where denotes the exponential map at , and denotes the matrix exponential (achieved through eigenvalue decomposition: ). Through Riemannian convolution, effective local feature aggregation and transformation are achieved on the manifold, further enhancing the geometric representation capability of the features.

3.3 Statistical feature enhancement module

The Statistical Feature Enhancement Module serves as the key bridge connecting geometric modeling and feature fusion in RGSTNet. Its core objective is to model the statistical distribution, perform adaptive reweighting, and optimize standardization of the manifold geometric features output by the Riemannian Geometry Module. By exploring the statistical regularities of the features, the module enhances the effective discriminative information and suppresses redundant noise, ultimately generating enhanced features that combine both geometric structure and statistical significance. These enhanced features provide high-quality inputs for the Manifold-Aware Transformer for cross-sample and cross-dimensional fusion.

3.3.1 Statistical analysis

Statistical analysis is the foundation for feature enhancement, aiming to comprehensively characterize the statistical distribution properties of geometric features in terms of mean, covariance, and probability density function (PDF). Given that the input features are located on manifolds such as the SPD manifold, statistical computation methods adapted to the manifold structure are required. For the manifold feature set output by the Riemannian Geometry Module, (where is the number of feature samples, and is the -th SPD manifold feature), the Fréchet mean on the manifold is first computed. This is defined as the point on the manifold that minimizes the sum of the squared geodesic distances from all samples:where denotes the geodesic distance on the SPD manifold, is the logarithmic map at point on the manifold, and is the Frobenius norm. This mean accurately reflects the center of the feature set on the manifold and avoids the problem where Euclidean mean would distort the manifold structure.

Based on the Fréchet mean, the covariance matrix of the manifold features is further constructed. First, each manifold feature is projected into the tangent space corresponding to using the logarithmic map. The tangent vector is:where is the dimension of the SPD matrix. After expanding the tangent vector, the covariance matrix is defined as:where is the Euclidean mean of the tangent vectors, is the regularization parameter, and is the identity matrix of dimension , ensuring the positive-definiteness of the covariance matrix. This matrix quantifies the linear correlation and distribution dispersion between features.

Assuming that the tangent vector set follows a multivariate normal distribution , the probability density function (PDF) for a single tangent vector is given by:where denotes the matrix determinant, and is the inverse of the covariance matrix, the PDF value reflects the statistical confidence of the feature: a larger indicates the feature is close to the statistical distribution center and is considered effective, while a smaller suggests the feature may be noise or an outlier.

3.3.2 Feature reweighting

Based on the PDF confidence obtained from statistical analysis, adaptive reweighting is applied to the original manifold features. By enhancing the contribution of high-confidence features and reducing the interference from low-confidence features, the discriminative power of the features is improved. The weighting process needs to consider the manifold structure to avoid distortion caused by Euclidean weighting: First, the PDF values are normalized to eliminate scale differences between samples, resulting in the normalized confidence :where and are the minimum and maximum PDF values across all samples, respectively. The normalized confidence ensures comparability of confidence.

Based on the normalized confidence, adaptive weights are designed. A temperature parameter is introduced to adjust the weight distinction, and the weight is computed as:where a smaller leads to a more significant weight difference between high- and low-confidence features. The weight set satisfies , ensuring consistency in the weighting process.

To adapt to the manifold structure, the weighted features are generated using a manifold-based weighted combination. The feature is first projected to the tangent space for weighting, and then mapped back to the manifold using the exponential map:where denotes the exponential map at . This operation not only adjusts the importance of features through weights but also fully preserves the manifold’s geometric structure.

To eliminate the scale differences between the reweighted features and enhance the model’s generalization ability and training stability, the reweighted features need to be standardized. Since the features remain on the SPD manifold, Riemannian normalization is used to ensure the standardized features have consistent manifold scale:where is the -dimensional identity matrix (the standard point on the manifold), and are the logarithmic and exponential maps at , and is a regularization parameter to avoid division by zero. The Riemannian norm ensures that the normalized features satisfy , keeping all features in the same manifold scale space and effectively avoiding the bias in model training caused by scale differences.

3.3.3 Enhanced features

After layers of optimization including statistical analysis, adaptive reweighting, and Riemannian normalization, the final enhanced feature set is obtained. This enhanced feature set retains the intrinsic manifold structure of the original geometric features while reinforcing high-confidence valid information, suppressing low-confidence noise interference, and eliminating the negative effects of scale differences. These enhanced features provide a strong foundation for the subsequent cross-sample and cross-dimensional attention fusion and classification prediction in the Manifold-Aware Transformer.

3.4 Manifold-aware transformer

As the core feature fusion and classification component of RGSTNet, the Manifold-Aware Transformer breaks through the limitations of traditional Transformers, which rely on Euclidean space, by adapting the attention mechanism, feed-forward network, and classification head to the Riemannian manifold structure. This enables deep feature fusion and accurate classification prediction across samples and dimensions. The module follows the logic chain of “manifold feature adaptation - attention aggregation - nonlinear transformation - classification output,” transforming the enhanced features output by the Statistical Feature Enhancement Module into classification features with global discriminative power, completing the closed loop from feature fusion to prediction output.

3.4.1 Riemannian attention

Riemannian Attention aims to capture the global dependencies between manifold features by introducing Riemannian geometric metrics to modify the attention weight computation, thus avoiding the destruction of manifold structure by Euclidean space distance metrics. Let the input enhanced feature set be (where is the number of feature samples and is the -th normalized SPD manifold feature). First, introduce a CLS token (initialized as the identity matrix on the manifold , where is the dimension of the SPD matrix), and construct the extended feature set , denoted as (where , and ).

To compute the attention weights, each manifold feature is first mapped to the query (Query), key (Key), and value (Value) spaces via a manifold-adapted linear projection. Define learnable manifold projection matrices , and perform the projection through Riemannian multiplication:where denotes matrix multiplication on the Riemannian manifold (achieved through eigenvalue decomposition: , with and being the eigen-decomposition of ), and represent the query, key, and value matrices of the -th feature.

The attention weight computation is based on manifold distance metrics, defining the Riemannian attention score between the query and key as:where is the geodesic distance on the SPD manifold (defined as before), and is the temperature coefficient used to adjust the smoothness of the attention score distribution.

Based on the attention weights, the value matrix is aggregated to obtain the attention output. Given the nonlinearity of manifold structures, Riemannian weighted averaging is used for aggregation:where is the local Fréchet mean of the value matrix set, and and are the logarithmic and exponential maps at . This aggregation method retains the manifold structure while performing global feature attention fusion.

To enhance the model’s expressiveness, a multi-head attention mechanism is introduced, dividing the query, key, and value spaces into heads. Each head independently computes the attention output, which is then fused via manifold concatenation and projection:where is the attention output of the -th head, denotes manifold feature concatenation (achieved by extending the feature dimension), and is the multi-head fusion projection matrix. The final output is .

3.4.2 Riemannian feed-forward network

The Riemannian Feed-Forward Network (RFFN) is used to perform nonlinear transformation and dimensional enhancement on the attention output features. By adapting two layers of nonlinear mappings to the manifold structure, the RFFN improves the nonlinear expressiveness of the features. The network consists of two core components: affine transform and MLP. The process is as follows:

First, an affine transform is applied to the attention output feature , performing translation and scaling on the manifold:where is the scaling coefficient, is the translation coefficient, is the learnable weight matrix, and is the learnable bias matrix (initialized as the identity matrix). The operation denotes scalar multiplication on the manifold (achieved by eigenvalue decomposition: ).

Next, an MLP is used for nonlinear transformation, with the first layer using a Riemannian ReLU activation function (Riemannian ReLU):where and are the eigen-decomposition results of , and is the regularization parameter to avoid matrix singularity after activation.

The second layer MLP performs dimensional reduction and feature refinement, outputting the final feed-forward network result:where is the learnable weight matrix and is the learnable bias matrix. The result is connected to the attention output feature via a residual connection (addition on the manifold: , where is the Fréchet mean of and ), and layer normalization (Riemannian layer normalization) is applied to obtain the final Transformer encoder output:where , and , are the mean and variance of the eigenvalues, and is the regularization parameter.

3.4.3 classification head

The Classification Head is responsible for transforming the CLS feature output by the Transformer encoder into classification prediction results. This is achieved through mapping from the manifold to Euclidean space, dimensional compression, and softmax classification. First, the CLS feature (corresponding to the final state of the initial CLS token) is extracted from the encoder output, and is projected into Euclidean space via the logarithmic map:where denotes the matrix logarithm, and is the symmetric matrix in Euclidean space.

Next, is flattened into a vector form (where denotes matrix vectorization), and is passed through a two-layer fully connected network (MLP) for dimensional compression:where , are learnable weight matrices ( is the hidden layer dimension, is the number of classes), and , are bias vectors, with as the activation function.

Finally, the class probability distribution is computed using the softmax function, producing the classification prediction results:where is the -th element of , and is the class probability vector, with representing the probability that the input sample belongs to class .

Through Riemannian Attention capturing global dependencies, Riemannian Feed-Forward Network enhancing nonlinear expression, and the Classification Head achieving precise prediction, the Manifold-Aware Transformer completes the deep fusion of manifold features and classification tasks, providing end-to-end image classification capabilities for RGSTNet.

3.5 Loss functions

To achieve end-to-end effective training of RGSTNet while ensuring feature discriminability, manifold structure consistency, and statistical distribution rationality, a joint loss function is designed, consisting of Label Smoothing Loss, Riemannian Contrastive Loss, and MMD + Fréchet Loss. Each loss function serves its purpose and optimizes collaboratively, with Label Smoothing Loss optimizing classification prediction accuracy, Riemannian Contrastive Loss enhancing the inter-class differentiation and intra-class aggregation of manifold features, and MMD + Fréchet Loss constraining the statistical distribution consistency of features. The global optimization objective is then formed through a weighted combination:where are the loss weight coefficients used to balance the contributions of each loss term, is the Label Smoothing Loss, is the Riemannian Contrastive Loss, and is the MMD + Fréchet Loss.

3.5.1 Label smoothing loss

The Label Smoothing Loss aims to alleviate the overfitting problem caused by traditional hard labels (one-hot encoding) by smoothing the true labels, guiding the model to learn more generalized feature representations and avoid excessive confidence in a single class. Let the true label of a training sample be (where is the number of classes), and the corresponding one-hot encoded label be (with only the -th position being 1 and the rest being 0). The smoothed label distribution is defined as:where is the smoothing coefficient, controlling the degree of label smoothing. Let the model’s classification prediction probability distribution be (which is the output of the Classification Head’s softmax, where represents the probability of the sample belonging to class ), then the Label Smoothing Loss is defined as:where is a small regularization term to avoid numerical instability caused by . This loss function introduces label uncertainty, reducing the model’s sensitivity to incorrect labels and encouraging the model to learn more robust classification features, thereby improving generalization.

3.5.2 Riemannian contrastive loss

Riemannian Contrastive Loss is designed for manifold features. It strengthens the inter-class separation and intra-class aggregation of features on the Riemannian manifold by reducing the manifold distance for same-class samples and increasing the manifold distance for different-class samples. Let the batch size be , and the manifold-enhanced feature for each sample be (where ), and the true class of sample be . Define the set of same-class samples as (positive sample set), and the set of different-class samples as (negative sample set).

To quantify the similarity between samples on the manifold, the Riemannian similarity is defined based on the geodesic distance:where is the geodesic distance on the SPD manifold, is the similarity adjustment parameter, and , with smaller distances leading to higher similarity.

Riemannian Contrastive Loss is defined as:where is the size of the positive sample set . This loss function, through logarithmic likelihood optimization, encourages the similarity of same-class samples to dominate the total similarity sum, thereby achieving intra-class manifold aggregation and inter-class manifold separation, enhancing feature discriminability.

3.5.3 MMD + Fréchet loss

MMD + Fréchet Loss is used to constrain the statistical distribution of the model’s output features to match the target distribution, such as the global feature distribution of the training set, thereby ensuring the statistical rationality of the features by quantifying the differences between the two distributions. This loss consists of MMD (Maximum Mean Discrepancy) and Fréchet distance, which constrain the distributions from the perspectives of kernel space distance and manifold center - covariance matching, respectively. Let the set of model output features be (batch features), and the target feature set be (global training set features, with being the number of global samples). Here, represents the dimension of the SPD matrix, and denotes the global Fréchet mean. First, the manifold features are projected into Euclidean space through logarithmic mapping, forming the tangent vector sets and , where denotes the matrix vectorization operation.

The MMD part calculates the distribution difference using a Gaussian kernel function, while the Fréchet part quantifies the distribution matching degree based on the mean and covariance. The final loss is obtained by the weighted sum of both parts:where is the Gaussian kernel bandwidth, is the kernel space, with and being the mean and covariance of , and and being the mean and covariance of . The denotes the matrix trace, and is the square root of the covariance matrix. The coefficient balances the contributions of the two parts. This loss ensures that the model output features have stable statistical distributions that align with global patterns, improving the model’s generalization ability and training stability.

4 Experiment

4.1 Datasets

To evaluate the effectiveness of the proposed RGSTNet framework, this study uses three widely adopted visual benchmark datasets: CIFAR-10, Fashion-MNIST, and MNIST. These datasets are commonly used to verify the representation learning and classification capabilities of visual models under different image types, resolutions, and semantic complexities. Although they are not collected directly from social networking platforms, they provide controlled and reproducible benchmark environments for evaluating the core visual content classification capability of the proposed method. Such capability is a fundamental component of intelligent social network applications, including visual content moderation, harmful content screening, misinformation-related media analysis, and automated platform governance. Detailed information about each dataset is summarized in Table 1.

TABLE 1

Dataset# Images# LabelsImage sizeType
CIFAR-1060,0001032 × 32RGB
FashionMNIST60,0001028 × 28Gray
MNIST60,0001028 × 28Gray

CIFAR-10, Fashion-MNIST, and MNIST datasets.

CIFAR-10 (): This dataset consists of 60,000 color RGB images with a resolution of 32 32 pixels, divided into 10 mutually exclusive classes (e.g., airplanes, cars, birds, cats). Each class contains exactly 6,000 images, with 50,000 images used for training and 10,000 for testing. The dataset features rich object diversity and natural scene variations, making it a standard benchmark for evaluating the robustness of image classification models to spatial details and color information.

Fashion-MNIST (): As a grayscale image dataset, it includes 60,000 samples with a resolution of 28 28 pixels, categorized into 10 fashion-related classes (e.g., T-shirts, trousers, dresses, sneakers). Similar to the split ratio of MNIST, 50,000 samples are allocated for training and 10,000 for testing. The dataset is designed to replace traditional MNIST for more realistic evaluation, as its samples have more complex texture and shape variations, posing greater challenges to feature extraction and discriminative learning.

MNIST (): A classic grayscale handwritten digit dataset, comprising 60,000 samples (50,000 for training and 10,000 for testing) with a resolution of 28 28 pixels. It covers 10 classes corresponding to digits 0–9. The dataset is characterized by simple background and clear digit contours, serving as a basic benchmark to verify the model’s fundamental classification capability and training stability.

In future work, we will further extend the evaluation to real-world social network datasets involving multimodal posts, user-generated images, misinformation-related media, and harmful visual content, so as to more comprehensively assess the practical deployment value of RGSTNet in social network governance scenarios.

4.2 Experimental setup

To ensure experimental reproducibility, computational fairness, and numerical stability, all experiments were conducted under a unified hardware and software environment. The implementation was based on PyTorch, and all compared models were trained and evaluated using the same data partition, preprocessing strategy, optimization protocol, and evaluation metrics. Although the current experiments are performed on standard visual benchmark datasets, the experimental design aims to verify the fundamental classification, distribution modeling, and non-Euclidean representation learning ability of the proposed RGSTNet. These capabilities are essential for subsequent intelligent simulation interpretation and early warning of risks in geological energy resources and underground spaces, where seismic images, simulation fields, monitoring maps, and engineering inspection images can be treated as structured visual or image-like data.

4.2.1 Hardware and software environment

All experiments were performed on a workstation equipped with a high-performance CPU and GPU to support large-scale matrix operations, eigenvalue decomposition, Riemannian logarithmic/exponential mappings, and Transformer-based feature fusion. The detailed hardware and software configurations are shown in Table 2. During training, GPU acceleration was used for convolution, attention computation, matrix multiplication, and batch-level feature optimization. CPU resources were mainly used for data loading, preprocessing, and auxiliary numerical computation. The CUDA and cuDNN versions were fixed to avoid performance variation caused by different low-level acceleration libraries.

TABLE 2

CategoryConfiguration details
Hardware EnvironmentCPU: Intel Core i9-13900K, 24 cores, 32 threads, max turbo frequency 5.8 GHz
GPU: NVIDIA GeForce RTX 4090, 24 GB GDDR6X VRAM
Memory: 64 GB DDR5, 5600 MHz
Storage: 2 TB NVMe SSD, PCIe 4.0
Power Supply: 1600 W, 80+ Titanium
Operating System: Ubuntu 22.04 LTS, 64-bit
Software EnvironmentDeep Learning Framework: PyTorch 2.1.0, TorchVision 0.16.0
Geometric Computing Library: PyTorch Geometric 2.4.0
Statistical Computing Library: NumPy 1.26.0, SciPy 1.11.3
Data Processing Library: Pandas 2.1.1, OpenCV 4.8.1
Visualization Tool: Matplotlib 3.8.0, Seaborn 0.13.0
Programming Language: Python 3.10.12
CUDA Toolkit: 12.2
cuDNN: 8.9.2

Hardware and software configuration.

4.2.2 Data preprocessing

For all datasets, images were first resized or padded to ensure consistent input dimensions within each dataset. CIFAR-10 images were kept at with three RGB channels, while Fashion-MNIST and MNIST images were kept at with one grayscale channel. Pixel values were normalized to the range [0, 1] and further standardized using the mean and standard deviation calculated from the training set. For RGB images, normalization was performed independently for each channel. For grayscale images, a single-channel normalization strategy was adopted.

To improve robustness and reduce overfitting, lightweight data augmentation was applied during training. For CIFAR-10, random horizontal flipping and random cropping with padding were used. For Fashion-MNIST and MNIST, random affine transformation with small rotation and translation ranges was adopted to simulate local deformation and observation uncertainty. No data augmentation was used during testing. The same preprocessing pipeline was applied to all baseline models and the proposed RGSTNet to ensure fair comparison.

4.2.3 Training and testing protocol

Each dataset was divided according to its official training and testing split. The model was trained only on the training set and evaluated on the corresponding test set. No test samples were used during training, hyperparameter tuning, or feature distribution estimation. For each experiment, the final reported results were obtained by averaging three independent runs with different random seeds. This strategy reduces the influence of random initialization and mini-batch sampling, thereby improving the statistical reliability of the reported performance.

In the effectiveness verification experiment, MNIST was used to evaluate the cross-dataset generalization ability of RGSTNet. The model was trained on CIFAR-10 and Fashion-MNIST without additional fine-tuning on MNIST. This setting was designed to test whether the learned geometric-statistical representation can transfer to unseen visual distributions. Such a protocol is meaningful for geological energy and underground space applications, where the distribution of monitoring or simulation data may change across different sites, geological structures, sensor systems, and engineering conditions.

4.2.4 Implementation of RGSTNet

The proposed RGSTNet was implemented as an end-to-end trainable framework consisting of the Riemannian Geometry Module, the Statistical Feature Enhancement Module, and the Manifold-Aware Transformer. For each input image, local pixels or image patches were regarded as nodes, and an adaptive -nearest-neighbor graph was constructed according to the mixed distance metric defined in the Methods section. The graph features were then mapped into the SPD manifold through local covariance estimation and positive-definite correction.

To ensure numerical stability in manifold computation, all SPD matrices were regularized by adding a small diagonal perturbation before eigenvalue decomposition. Eigenvalues smaller than the threshold were clipped to avoid singular matrices. Matrix logarithm and matrix exponential operations were computed through eigenvalue decomposition. During backpropagation, gradients were propagated through differentiable matrix operations provided by PyTorch. The statistical feature enhancement module estimated the Fréchet mean, covariance matrix, and probability density values in the tangent space. These statistical quantities were then used for distribution-aware feature reweighting and Riemannian normalization.

The Manifold-Aware Transformer used manifold-adapted query, key, and value projections. Attention scores were calculated based on geodesic distances instead of conventional Euclidean dot-product similarity. The final CLS feature was mapped from the SPD manifold to Euclidean space through logarithmic mapping and then fed into the classification head. This design enables RGSTNet to jointly exploit local geometric structures, global dependencies, and statistical distribution information.

4.2.5 Optimization strategy

All models were optimized using the AdamW optimizer. The initial learning rate was set to , and the minimum learning rate was set to . A cosine annealing learning-rate scheduler was adopted to gradually reduce the learning rate during training. Weight decay was applied to reduce overfitting. The total number of training epochs was set to 200, and the batch size was set to 128. The same optimizer and training epochs were used for all compared methods unless otherwise specified.

For RGSTNet, the total training objective consisted of three parts: Label Smoothing Loss, Riemannian Contrastive Loss, and MMD+Fréchet Loss. The label smoothing term improves classification robustness by reducing over-confident predictions. The Riemannian contrastive term enhances intra-class compactness and inter-class separability on the manifold. The MMD+Fréchet term constrains the statistical distribution of learned features and improves robustness under distribution shifts. The final loss function was optimized as:

The loss weights were empirically set according to validation stability and are reported in Table 3.

TABLE 3

Hyperparameter categoryParameter nameValue
Training configurationBatch Size128
Initial Learning Rate
Minimum Learning Rate
Weight Decay
Training Epochs200
OptimizerAdamW
Learning Rate SchedulerCosine Annealing
Number of Independent Runs3
Optimizer parameters0.9
0.999
Graph construction parametersNumber of Nearest Neighbors 8
Distance Weighting Factor 0.5
Gaussian Kernel Bandwidth 1.0
Manifold and numerical parametersFeature Reweighting Temperature 0.1
Attention Score Temperature 64
Regularization
Regularization
Regularization
Eigenvalue Threshold
Transformer parametersNumber of Attention Heads 4
Hidden Dimension 256
Dropout Rate0.1
CLS Token InitializationIdentity Matrix
Loss function weights for Label Smoothing Loss1.0
for Riemannian Contrastive Loss2.0
for MMD+Fréchet Loss1.5
Kernel and statistical parametersGaussian Kernel Bandwidth 0.5
Similarity Adjustment Parameter 1.0
Label Smoothing Coefficient0.1
MMD-Fréchet Balance Coefficient0.5

Hyperparameter settings for model training.

4.2.6 Hyperparameter settings

The key hyperparameters used in RGSTNet are summarized in Table 3. These parameters include general training settings, optimizer parameters, manifold regularization coefficients, statistical distribution parameters, and loss function weights. To ensure fairness, the main training parameters, such as batch size, learning rate, optimizer, and training epochs, were kept consistent across all compared models.

4.2.7 Baseline implementation

To provide a comprehensive comparison, RGSTNet was compared with representative CNN-based, lightweight, graph-related, state-space, and Transformer-based models, including ResNet, EfficientNet, DenseNet, ConvNeXt, MobileNetV2, MambaNet, ViT, Swin Transformer, DeiT, T2T-ViT, and CaiT. For fairness, all baseline models were trained using the same training/test split, preprocessing strategy, optimizer type, batch size, and number of epochs. For models with publicly available standard implementations, the official or widely used PyTorch implementations were adopted. The final classification layer of each baseline model was adjusted according to the number of classes in each dataset.

4.2.8 Evaluation protocol

The model performance was evaluated using Precision, Recall, Accuracy, and F1-Score. These metrics were computed on the test set after training was completed. Accuracy measures the overall proportion of correctly classified samples. Precision reflects the reliability of positive predictions, Recall measures the ability to identify target samples, and F1-Score provides a balanced evaluation of Precision and Recall. For multi-class classification, macro-averaged metrics were used to reduce bias caused by class-frequency differences. In addition, confusion matrices were generated to analyze category-level classification behavior and inter-class confusion patterns.

4.2.9 Reproducibility settings

To improve reproducibility, random seeds were fixed for Python, NumPy, and PyTorch in each independent run. The data loading order, model initialization, and augmentation randomness were controlled by the corresponding seed. All experiments were repeated three times, and the average results were reported. The same computational environment and hyperparameter configuration were used throughout the experiments.

4.3 Quantitative comparison results

4.3.1 Performance on general image datasets

Table 4 presents a detailed quantitative comparison between the proposed RGSTNet and 18 mainstream image classification models on the CIFAR-10 and Fashion-MNIST datasets, using four core evaluation metrics: Precision, Recall, Accuracy, and F1-Score.

TABLE 4

ModelCIFAR-10FashionMNIST
PrecisionRecallAccuracyF1-scorePrecisionRecallAccuracyF1-score
ResNet-18 ()0.910.920.910.910.870.880.880.87
EfficientNet-B0 ()0.930.920.920.920.890.900.890.89
DenseNet ()0.900.910.900.900.860.870.860.86
ViT ()0.940.930.940.940.900.910.910.90
ConvNeXt ()0.920.910.910.910.880.890.880.88
MambaNet ()0.890.900.890.890.850.860.850.85
SqueezeNet ()0.880.890.880.880.840.850.840.84
AlexNet ()0.900.910.900.900.850.860.850.85
LeNet-5 ()0.870.880.870.870.800.810.800.80
VGG-16 ()0.920.910.920.920.870.880.870.87
MobileNetV2 ()0.910.900.900.900.850.860.850.85
Xception ()0.930.920.930.930.890.900.890.89
ResNet-50 ()0.940.930.940.940.900.910.900.90
NASNet-A ()0.950.940.950.950.910.920.910.91
Swin Transformer ()0.950.940.950.950.920.930.920.92
DeiT ()0.940.930.940.940.910.920.910.91
T2T-ViT ()0.950.940.950.950.910.920.910.91
CaiT ()0.960.950.960.960.930.940.930.94
RGSTNet (Ours)0.980.970.980.980.950.960.950.96

Comparison of different methods on the CIFAR-10 and fashionMNIST datasets.

Bold values indicate the best result for each evaluation metric.

On the CIFAR-10 dataset, RGSTNet achieves Precision, Recall, Accuracy, and F1-Score values of 0.98, 0.97, 0.98, and 0.98, respectively, demonstrating clear advantages over existing state-of-the-art methods. Compared with CaiT, the previously best-performing model, RGSTNet improves Precision, Recall, Accuracy, and F1-Score by 2 percentage points each. It also outperforms Swin Transformer and NASNet-A, both of which achieve an Accuracy of 0.95, by 3 percentage points in Accuracy. In comparison with classical models such as ResNet-50 and ViT, both with an Accuracy of 0.94, RGSTNet achieves a 4-percentage-point improvement in Accuracy.

On the Fashion-MNIST dataset, RGSTNet attains Precision, Recall, Accuracy, and F1-Score values of 0.95, 0.96, 0.95, and 0.96, respectively. Compared with the leading baseline CaiT, which achieves an Accuracy of 0.93, RGSTNet improves Precision, Recall, Accuracy, and F1-Score by 2 percentage points each. In addition, it surpasses Swin Transformer, which achieves an Accuracy of 0.92, by 3 percentage points in Accuracy. It also exceeds strong conventional baselines such as ResNet-50, with an Accuracy of 0.90, by 5 percentage points and EfficientNet-B0, with an Accuracy of 0.89, by 6 percentage points in Accuracy.

These results demonstrate that the deep integration of Riemannian geometric modeling, statistical feature optimization, and manifold-aware global feature fusion in RGSTNet effectively alleviates the limitations of conventional Euclidean-based models and yields substantial improvements across key classification metrics.

4.3.2 Performance on geological dataset

To further validate the effectiveness of RGSTNet in real-world geological energy and underground space applications, we compare it with seven state-of-the-art methods (2024–2026) on the Netherlands F3 seismic facies classification dataset, as shown in Table 5. The compared baselines include TransNeXt (), MambaVision (), StarNet (), RMT (), SHViT (), , and , all representing the most recent advances in vision Transformers, hybrid architectures, and geometry-aware models.

TABLE 5

MethodYearPrecisionRecallAccuracyF1-score
TransNeXt ()202488.4288.0588.3188.23
MambaVision ()202489.1088.7488.9788.92
StarNet ()202487.9587.6087.8387.77
RMT ()202489.3588.9889.2289.16
SHViT ()202488.2087.8588.0888.02
202590.0589.7089.9389.87
202690.6290.2890.5190.45
RGSTNet (Ours)202692.3792.0492.2592.20

Comparison with state-of-the-art methods on the Netherlands F3 seismic facies dataset. Best results are in bold.

On the Netherlands F3 dataset, RGSTNet achieves Precision, Recall, Accuracy, and F1-Score of 92.37%, 92.04%, 92.25%, and 92.20%, respectively, outperforming all compared methods. Compared with the previous best-performing method , which achieves an Accuracy of 90.51%, RGSTNet improves Precision, Recall, Accuracy, and F1-Score by approximately 1.75 percentage points. Compared with , which achieves an Accuracy of 89.93%, RGSTNet improves Accuracy by 2.32 percentage points. RGSTNet also significantly outperforms earlier 2024 methods such as RMT (89.22% Accuracy) and MambaVision (88.97% Accuracy), with improvements of 3.03 and 3.28 percentage points in Accuracy, respectively.

These results demonstrate that RGSTNet’s Riemannian geometric modeling, statistical feature enhancement, and manifold-aware Transformer are particularly effective for geological datasets, where complex non-Euclidean structures (e.g., curved horizons, faults), heterogeneous distributions (e.g., noise, amplitude variations), and multi-scale spatial dependencies are prevalent. The performance gain over recent geometry-aware baselines (GeoFormer, RiemannFormer) further validates the effectiveness of the proposed distribution-aware Riemannian framework for intelligent simulation interpretation and risk early warning in geological energy and underground spaces.

4.3.3 Visual comparison across datasets

As shown in Figure 2, the significant advantages of RGSTNet in terms of Precision, Recall, Accuracy, and F1-Score are clearly demonstrated. Its performance exceeds that of comparison models such as the ResNet series, ViT, Swin Transformer, and CaiT across all metrics, with particularly remarkable performance in the Precision metric on the CIFAR-10 dataset, creating a noticeable performance gap. We further validated the effectiveness of the RGSTNet design, which integrates Riemannian geometric modeling, statistical feature optimization, and the Manifold-Aware Transformer, through this visualization result. It proves that RGSTNet is able to fully leverage the complementary value of geometric structure and statistical features when handling non-Euclidean space image data, thus surpassing traditional models in key classification metrics, providing intuitive visual support for subsequent model performance analysis and advantage demonstration.

FIGURE 2

4.4 Ablation experiment

To verify the effectiveness of each core module in the RGSTNet framework and their synergistic effects, ablation experiments are conducted on the CIFAR-10 dataset, with the baseline model set as ViT. The experimental results are summarized in Table 6, focusing on Recall, Accuracy, and F1-Score. When only the Riemannian Geometry Module (RGM) is added to the baseline, the model’s Recall, Accuracy, and F1-Score reach 0.94, 0.95, and 0.95 respectively, showing an improvement of one to two percentage points compared to the baseline (0.93, 0.94, 0.94), which confirms that capturing the intrinsic geometric structure of data through RGM effectively enhances feature representation. Adding only the Statistical Feature Enhancement Module (SFEM) to the baseline yields metrics of 0.95, 0.95, and 0.95, indicating that optimizing feature statistical distribution via SFEM strengthens feature discriminability. The independent addition of the Manifold-Aware Transformer (MAT) results in metrics of 0.95, 0.94, and 0.94, demonstrating that MAT’s manifold-adapted global feature fusion alleviates the feature distortion problem of traditional Transformers. When combining RGM and SFEM, the model’s Accuracy further improves to 0.96, reflecting the complementary advantages of geometric structure mining and statistical feature optimization. The combination of RGM and MAT maintains stable performance with metrics of 0.94, 0.95, and 0.95. Notably, when all three modules (RGM + SFEM + MAT) are integrated into the baseline, the model achieves the best performance: Recall of 0.97, Accuracy of 0.98, and F1-Score of 0.98, which are 4, 4, and 4 percentage points higher than the baseline, respectively.

TABLE 6

Model variantsModulesMetrics
RGMSFEMMATRecallAccuracyF1-score
Baseline (ViT)XXX0.930.940.94
Baseline + RGMXX0.940.950.95
Baseline + SFEMXX0.950.950.95
Baseline + MATXX0.950.940.94
Baseline + RGM + SFEMX0.950.960.94
Baseline + RGM + MATX0.940.950.95
Baseline+ RGM + SFEM + MAT0.970.980.98

Ablation experiments on the CIFAR-10 dataset.

To verify the effectiveness of each component in the proposed joint loss function and their synergistic optimization effect, additional ablation experiments are conducted on the CIFAR-10 dataset, with the full RGSTNet (integrating RGM, SFEM, and MAT) as the base model. The experimental results are summarized in Table 7, focusing on Recall, Accuracy, and F1-Score. When only using a single loss function, the model performance is relatively limited: employing Label Smoothing Loss alone yields Recall of 0.90, Accuracy of 0.91, and F1-Score of 0.90; using Riemannian Contrastive Loss independently achieves metrics of 0.91, 0.90, and 0.91; and adopting MMD+Fréchet Loss solely results in 0.90, 0.91, and 0.90. When combining and , the model performance is significantly improved, with Recall, Accuracy, and F1-Score reaching 0.96, 0.96, and 0.97 respectively—an increase of five to six percentage points compared to single loss functions. This indicates that the combination of classification accuracy optimization and inter-class/intra-class feature constraint forms a strong complementary effect. Notably, when integrating all three loss functions , the model achieves the optimal performance: Recall of 0.97, Accuracy of 0.98, and F1-Score of 0.98. Compared to the dual-loss combination, this represents an additional 1 percentage point improvement in key metrics, confirming that effectively constrains the statistical distribution consistency of features, further enhancing the model’s generalization ability and feature discriminability. These results fully demonstrate that the joint loss function designed in this paper—through the synergistic effect of , , and comprehensively optimizes classification accuracy, geometric structure consistency, and statistical distribution rationality, laying a key foundation for the superior performance of RGSTNet.

TABLE 7

Model variantsModulesMetrics
RecallAccuracyF1-score
XX0.900.910.90
XX0.910.900.91
XX0.900.910.90
+ X0.960.960.97
+ + 0.970.980.98

Ablation experiments on loss functions.

As shown in Figure 3, the impact of different loss function configurations on model performance shows significant differences. Figure a clearly demonstrates the trend in Recall, where single loss functions correspond to lower Recall levels, while the complete configuration integrating all three loss functions significantly improves Recall, highlighting the strengthening effect of multi-objective collaborative optimization on the model’s ability to identify positive samples. Figure b shows that Accuracy and F1-Score increase simultaneously as the loss function combination is refined. Under the complete loss function configuration, both metrics reach their optimal values while maintaining good consistency, verifying the synergistic effect of the joint loss function in optimizing classification accuracy, geometric structure consistency, and statistical distribution rationality.

FIGURE 3

4.5 Effectiveness verification

To further validate the generalization ability of the proposed RGSTNet framework, we conduct cross-dataset verification on the MNIST dataset without any additional training or parameter fine-tuning (the model is only trained on CIFAR-10 and Fashion-MNIST). As shown in Table 8, RGSTNet (denoted as RGSTNet (Ours)) achieves the best overall performance among compared models with Precision of 0.99, Recall of 0.98, Accuracy of 0.96, and F1-Score of 0.97—outperforming ResNet-18 (Accuracy 0.93, Precision 0.95) by 3 percentage points in Accuracy and 4 percentage points in Precision, MambaNet (Accuracy 0.95, Precision 0.97) by 1 percentage point in Accuracy and 2 percentage points in Precision, and maintaining comparable Accuracy and F1-Score to ViT (Accuracy 0.96, F1-Score 0.97) while achieving 1 percentage point higher Precision. These results fully demonstrate that RGSTNet, relying on the deep integration of Riemannian geometric modeling, statistical feature optimization, and manifold-aware global fusion, can learn robust and generalizable feature representations from diverse training data, effectively adapting to unseen simple-structured grayscale handwritten digit data, thus verifying its strong cross-dataset generalization ability.

TABLE 8

MethodPrecisionRecallAccuracyF1-score
ResNet-180.950.940.930.94
MambaNet0.970.960.950.96
ViT0.980.970.960.97
RGSTNet (Ours)0.990.980.960.97

Comparison of different methods on the MNIST datasets.

Bold values indicate the best result for each evaluation metric.

4.6 Confusion matrix

Figures 4a,b present the confusion matrices of RGSTNet on the CIFAR-10 and Fashion-MNIST datasets, respectively, providing a category-level view of classification accuracy and inter-class prediction confusion. These visualizations intuitively reflect the model’s classification behavior across different categories and reveal the degree of confusion between visually similar classes.

FIGURE 4

On the CIFAR-10 dataset, RGSTNet achieves consistently high and well-balanced classification performance across all categories, with only limited confusion observed among a few classes sharing similar visual characteristics. Most categories are correctly identified, indicating that the model effectively captures the intrinsic geometric structure and statistical characteristics of different object classes.

On the Fashion-MNIST dataset, the model also maintains strong classification performance across all categories. Although slight overlap appears among some clothing categories due to similarities in texture and contour, the overall degree of confusion remains low and does not materially affect the model’s overall classification performance.

The confusion matrix results on both datasets further verify the effectiveness of RGSTNet, whose design combines Riemannian geometric modeling for intrinsic structure exploration with statistical feature optimization for improved feature discriminability. The low level of inter-class confusion in category differentiation also suggests that the model has strong potential for stable deployment in real-world applications.

5 Limitations

Although RGSTNet achieves strong performance on both general image benchmarks and the Netherlands F3 geological dataset, several limitations remain that warrant further investigation.

First, the geological validation in this study relies primarily on the Netherlands F3 seismic facies dataset. While this dataset is representative of subsurface stratigraphic interpretation tasks, it does not fully capture the diversity of geological energy and underground space scenarios. Broader validation on datasets such as reservoir characterization, geothermal field monitoring, tunnel deformation detection, and underground gas storage risk assessment is needed to establish the generalizability of the proposed method across heterogeneous subsurface conditions.

Second, the current model is designed for single-modality, image-type data. In practical geological energy development and underground engineering, decision-making typically depends on multi-source heterogeneous information, including well-logging curves, geophysical inversion results, numerical simulation outputs, and real-time IoT sensor monitoring. RGSTNet does not yet incorporate such multi-modal fusion, which may limit its applicability in integrated risk evaluation and early warning systems.

Third, the Riemannian geometric modeling and manifold-aware attention introduce additional computational overhead compared with conventional Euclidean architectures. Although this cost is acceptable for offline analysis and simulation, it may pose challenges for real-time, resource-constrained deployment scenarios such as edge-based underground monitoring. The trade-off between geometric expressiveness and computational efficiency requires further optimization.

Fourth, the present work formulates risk-related tasks as classification problems. Real-world geological risk assessment and early warning often involve temporal dynamics, spatial continuity, and uncertainty quantification that go beyond static classification. Extending RGSTNet to spatiotemporal modeling and probabilistic risk estimation remains an open direction.

Addressing these limitations through multi-source data fusion, lightweight manifold computation, spatiotemporal extension, and broader domain validation will be the focus of our future work, with the goal of developing a more comprehensive and deployable framework for intelligent simulation and safety early warning in geological energy resources and underground spaces.

6 Conclusion

This paper proposes RGSTNet, a Distribution-Aware Riemannian Transformer for intelligent simulation-oriented classification and early warning of risks in geological energy resources and underground spaces. The proposed method addresses the limitations of conventional Euclidean deep learning models by integrating Riemannian geometric representation learning, statistical feature enhancement, and manifold-aware global attention fusion. Specifically, RGSTNet consists of three core components: a Riemannian Geometry Module for manifold-based structural modeling of geological and underground spatial data, a Statistical Feature Enhancement Module for distribution-aware feature optimization under complex heterogeneous conditions, and a Manifold-Aware Transformer for global feature fusion under Riemannian metrics. In addition, a joint loss function combining label smoothing, Riemannian contrastive learning, and MMD-Fréchet distribution alignment is designed to improve classification accuracy, geometric consistency, and statistical robustness.

Experimental results on CIFAR-10, Fashion-MNIST, and the Netherlands F3 seismic facies dataset demonstrate that RGSTNet achieves superior performance compared with representative CNN-, GNN-, and Transformer-based baselines in terms of Precision, Recall, Accuracy, and F1-score. Notably, on the Netherlands F3 geological dataset, RGSTNet outperforms state-of-the-art methods by substantial margins, validating its effectiveness in capturing non-Euclidean structural dependencies and heterogeneous distribution characteristics inherent in subsurface data. The ablation studies further verify the effectiveness of each proposed module and the joint optimization objective. These results indicate that RGSTNet provides a promising technical foundation for intelligent geological energy resource evaluation, underground space safety assessment, simulation-driven risk classification, and early warning decision support.

Future work will focus on three directions. First, we will extend RGSTNet to a broader range of real-world geological energy and underground space applications, including reservoir heterogeneity characterization, fracture network identification, geothermal anomaly detection, and underground infrastructure risk assessment. Second, we will incorporate multi-source heterogeneous data fusion, jointly modeling seismic images, well-logging curves, numerical simulation fields, and real-time monitoring signals to support integrated risk evaluation and early warning. Third, we will explore lightweight Riemannian operations, sparse manifold attention, and edge-cloud collaborative deployment strategies to improve the scalability and real-time performance of RGSTNet in practical geological energy and underground space monitoring systems. The remaining challenges discussed in the Limitations section will guide these efforts toward a more comprehensive and deployable framework.

Statements

Data availability statement

The original contributions presented in the study are included in the article/Supplementary Material, further inquiries can be directed to the corresponding author.

Author contributions

ZG: Conceptualization, Data curation, Formal Analysis, Investigation, Methodology, Project administration, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review and editing.

Funding

The author(s) declared that financial support was not received for this work and/or its publication.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

References

  • 1

    BronsteinM. M.BrunaJ.LeCunY.SzlamA.VandergheynstP. (2017). Geometric deep learning: going beyond euclidean data. IEEE Signal Process. Mag.34 (4), 1842. 10.1109/MSP.2017.2698990

  • 2

    CholletF. (2017). “Xception: deep learning with depthwise separable convolutions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 12511258. 10.1109/CVPR.2017.195

  • 3

    DosovitskiyA.BeyerL.KolesnikovA.WeissenbornD.ZhaiX.UnterthinerT.et al (2021). “An image is worth 16× 16 words: transformers for image recognition at scale,” in International Conference on Learning Representations (ICLR). 10.48550/arXiv.201.11929

  • 4

    DosovitskiyA. (2020). An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv preprint arXiv:201.11929. 10.48550/arXiv.2010.11929

  • 5

    FanQ.HuangH.ChenM.LiuH.HeR. (2024). “RMT: retentive networks meet vision transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 10.48550/arXiv.2309.11523

  • 6

    FonioS.EspositoR.AldinucciM. (2025). “Hyperbolic prototypical entailment cones for image classification,” in Proceedings of The 28th International Conference on Artificial Intelligence and Statistics, 33583366. 10.48550/arXiv.2304.08428

  • 7

    GaneaO.BécigneulG.HofmannT. (2018). Hyperbolic neural networks. Adv. Neural Information Processing Systems31, 53505360. 10.48550/arXiv.1805.09112

  • 8

    GrahamB.El-NoubyA.TouvronH.StockP.JoulinA.JégouH.et al (2021). “Levit: a vision transformer in convnet’s clothing for faster inference,” in Proceedings of the IEEE/CVF international conference on computer vision, 1225912269. 10.48550/arXiv.2104.01136

  • 9

    HanK.WangY.ChenH.ChenX.GuoJ.LiuZ.et al (2022). A survey on vision transformer. IEEE Trans. Pattern Analysis Mach. Intell.45 (1), 87110. 10.1109/TPAMI.2022.3151672

  • 10

    HatamizadehA.KautzJ. (2024). MambaVision: a hybrid mamba-transformer vision backbone, arXiv preprint arXiv:2407.08083. 10.48550/arXiv.2407.08083

  • 11

    HeK.ZhangX.RenS.SunJ. (2016). “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 770778. 10.1109/CVPR.2016.90

  • 12

    HowardA.SandlerM.ChuG.ChenL.-C.ChenB.TanM.et al (2019). “Searching for mobilenetv3,” in Proceedings of the IEEE/CVF international conference on computer vision, 13141324. 10.48550/arXiv.1905.02244

  • 13

    HuangG.LiuZ.Van Der MaatenL.WeinbergerK. Q. (2017). “Densely connected convolutional networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 47004708. 10.1109/CVPR.2017.243

  • 14

    IandolaF. N.HanS.MoskewiczM. W.AshrafK.DallyW. J.KeutzerK. (2016). SqueezeNet: alexnet-level accuracy with 50× fewer parameters and <0.5MB model size, arXiv preprint arXiv:1602.07360. 10.48550/arXiv.1602.07360

  • 15

    JiZ. (2025). RiemannFormer: a framework for attention in curved spaces. arXiv [Preprint], arXiv:2506.07405. 10.48550/arXiv.2506.07405

  • 16

    KriukB.GillS. K.AslamS.FakhrutdinovA. (2025). GFT: gradient focal transformer, arXiv preprint arXiv:2504.09852. 10.48550/arXiv.2504.09852

  • 17

    KrizhevskyA. (2009). Learning Multiple Layers of Features from Tiny Images. Technical Report. Toronto, ON, Canada: Department of Computer Science, University of Toronto.

  • 18

    KrizhevskyA.SutskeverI.HintonG. E. (2012). ImageNet classification with deep convolutional neural networks. Adv. Neural Inf. Process. Syst. (NeurIPS)60, 10971105. 10.1145/3065386

  • 19

    KuiX.JiangS.LiQ.PengY.HuZ.ZouB. (2025). GL-MambaNet: a global-local hybrid mamba network for medical image segmentation. Neurocomputing626, 129580. 10.1016/j.neucom.2025.01.022

  • 20

    LeCunY.BottouL.BengioY.HaffnerP. (1998). Gradient-based learning applied to document recognition. Proc. IEEE86 (11), 22782324. 10.1109/5.726791

  • 21

    LinT.ZhaH. (2008). Riemannian manifold learning. IEEE Transactions Pattern Analysis Machine Intelligence30 (5), 796809. 10.1109/TPAMI.2007.70735

  • 22

    LiuZ.LinY.CaoY.HuH.WeiY.ZhangZ.et al (2021a). “Swin transformer: hierarchical vision transformer using shifted window,” in Proceedings of the IEEE/CVF international conference on computer vision, 1001210022. 10.48550/arXiv.210.14030

  • 23

    LiuZ.HuH.LinY.YaoZ.LiZ.ShiY.et al (2021b). “Swin transformer: hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 1001210022. 10.48550/arXiv.2103.14030

  • 24

    LiuZ.MaoH.WuC.-Y.FeichtenhoferC.DarrellT.XieS. (2022a). “A convnet for the 2020s,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 1197611986. 10.48550/arXiv.2201.03545

  • 25

    LiuZ.MaoH.WuC.-Y.FeichtenhoferC.DarrellT.XieS. (2022b). “A ConvNet for the 2020s,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 1197611986. 10.1109/CVPR46437.2022.01172

  • 26

    LohitS.TuragaP. (2017). “Learning invariant Riemannian geometric representations using deep nets,” in Proceedings of the IEEE International Conference on Computer Vision Workshops, 13291338. 10.48550/arXiv.1708.09485

  • 27

    MaX.DaiX.BaiY.WangY.FuY. (2024). “Rewrite the stars,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 10.48550/arXiv.2403.19967

  • 28

    MasciJ.BoscainiD.BronsteinM.VandergheynstP. (2025). “Geodesic convolutional neural networks on Riemannian manifolds,” in Proceedings of the IEEE International Conference on Computer Vision Workshops, 3745. 10.48550/arXiv.1501.06297

  • 29

    MettesP.Ghadimi AtighM.Keller-ResselM.GuJ.YeungS. (2024). Hyperbolic deep learning in computer vision: a survey. Int. J. Comput. Vis.132 (9), 34843508. 10.1007/s11263-024-01492-3

  • 30

    MunirM.RahmanM. M.MarculescuR. (2025). “RapidNet: multi-level dilated convolution based Mobile backbone,” in IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 83028312. 10.48550/arXiv.2412.10995

  • 31

    NingE.LiW.FangJ.YuanJ.DuanQ.WangG. (2025). 3D-Guided multi-feature semantic enhancement network for person Re-ID. Inf. Fusion117, 102863. 10.1016/j.inffus.2024.102863

  • 32

    NingE.GaoW.LiuY.YuanJ.ZhangH.HeS.et al (2026a). Causally invariant video anomaly detection via counterfactual reasoning and prototype intervention. Pattern Recognit.179, 113809. 10.1016/j.patcog.2026.113809

  • 33

    NingE.WuL.XieS.YangJ.HuX.HuZ.et al (2026b). Disentangling identity from appearance: a semantic-hierarchical multi-level fusion framework for cloth-invariant person Re-Identification. Inf. Fusion133, 104346. 10.1016/j.inffus.2026.104346

  • 34

    NingE.MiaoJ.XieS.MaH.NingX. (2026c). Occluded person Re-Identification in multi-scenarios: a synergistic interaction framework with perception-aware optimization. Eng. Appl. Artif. Intell.166, 113674. 10.1016/j.engappai.2025.113674

  • 35

    OquabM.DarcetT.MoutakanniT.VoH.SzafraniecM.KhalidovV.et al (2023). Dinov2: Learning Robust Visual Features without Supervision. arXiv preprint arXiv:2304.07193. 10.48550/arXiv.2304.07193

  • 36

    ParikhD.Fein-AshleyJ.YeT.KannanR.PrasannaV. (2025). ClusterViG: Efficient Globally Aware Vision GNNs via Image Partitioning. arXiv preprint arXiv:2501.10640. 10.48550/arXiv.2501.10640

  • 37

    RadfordA.NarasimhanK.SalimansT.SutskeverI. (2018). Improving language Understanding by Generative Pre-training. San Francisco, CA, USA.

  • 38

    RadfordA.KimJ. W.HallacyC.RameshA.GohG.AgarwalS.et al (2021). “Learning transferable visual models from natural language supervision,” in International conference on machine learning, 87488763. 10.48550/arXiv.2103.00020

  • 39

    SaidS.HajriH.BombrunL.VemuriB. C. (2017). Gaussian distributions on Riemannian symmetric spaces: statistical learning with structured covariance matrices. IEEE Trans. Inf. Theory64 (2), 752772. 10.1109/TIT.2017.2788404

  • 40

    SandlerM.HowardA.ZhuM.ZhmoginovA.ChenL. C. (2018). “MobileNetV2: inverted residuals and linear bottlenecks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 45104520. 10.1109/CVPR.2018.00474

  • 41

    ShiD. (2024). “TransNeXt: robust foveal visual perception for vision transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 10.48550/arXiv.2311.17132

  • 42

    SimonyanK.ZissermanA. (2015). “Very deep convolutional networks for large-scale image recognition,” in International Conference on Learning Representations (ICLR). 10.48550/arXiv.1409.1556

  • 43

    TanM.LeQ. V. (2019). “EfficientNet: rethinking model scaling for convolutional neural networks,” in Proceedings of the 36th International Conference on Machine Learning (ICML), 61056114. 10.5555/3454287.3454311

  • 44

    TanM.LeQ. V. (2021). Efficientnetv2: smaller models and faster training, arXiv preprint arXiv:2104.00298, 5. 10.48550/arXiv.2104.00298

  • 45

    TouvronH.CordM.SablayrollesA.SynnaeveG.JégouH. (2021a). “Training data-efficient image transformers and distillation through attention,” in International Conference on Machine Learning (ICML), 1034710357. 10.48550/arXiv.2012.12877

  • 46

    TouvronH.BojanowskiM.LouisP. J. Y.CordA.SynnaeveG.JégouH. (2021b). “CaiT: class-attention in image transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 10.48550/arXiv.2106.14881

  • 47

    TuragaP.VeeraraghavanA.SrivastavaA.ChellappaR. (2011). Statistical computations on grassmann and stiefel manifolds for image and video-based recognition. IEEE Trans. Pattern Analysis Mach. Intell.33 (11), 22732286. 10.1109/TPAMI.2010.235

  • 48

    WuZ.PanS.ChenF.LongG.ZhangC.YuP. S. (2020). A comprehensive survey on graph neural networks. IEEE Transactions Neural Networks Learning Systems32 (1), 424. 10.1109/TNNLS.2020.2978386

  • 49

    WuH.XiaoB.CodellaN.LiuM.DaiX.YuanL.et al (2021). “CVT: introducing convolutions to vision transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2231. 10.48550/arXiv.2103.15808

  • 50

    XiaoH.RasulK.VollgrafR. (2017). Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms,arXiv preprint arXiv:1708.07747. 10.48550/arXiv.1708.07747

  • 51

    YuanL.ChenY.WangT.YuW.ShiY.JiangZ.et al (2021). Tokens-to-Token ViT: training vision transformers from scratch on ImageNet, arXiv preprint arXiv:2101.11986. 10.48550/arXiv.2101.11986

  • 52

    YunS.RoY. (2024). “SHViT: single-head vision transformer with memory efficient macro design,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 10.48550/arXiv.2401.16456

  • 53

    YunS.JeongM.KimR.KangJ.KimH. J. (2019). Graph transformer networks. Adv. Neural Inf. Process. Syst.32, 1196011970. 10.48550/arXiv.1911.06455

  • 54

    ZaheerM.KotturS.RavanbakhshS.PoczosB.SalakhutdinovR. R.SmolaA. J. (2017). Deep sets. Adv. Neural Information Processing Systems30, 33913401. 10.48550/arXiv.1703.06114

  • 55

    ZhangS.TongH.XuJ.MaciejewskiR. (2019). Graph convolutional networks: a comprehensive review. Comput. Soc. Netw.6 (1), 123. 10.1007/s00462-019-00205-0

  • 56

    ZhangX.CaoS.YuZ.WuZ.ZhangX.BaiX.et al (2025). GeoFormer: boosting object distinguishing and prompt understanding for cross-view object geo-localization. IEEE Trans. Geosci. Remote Sens.63, 116. 10.1109/TGRS.2025.3638946

  • 57

    ZheX.ChenS.YanH. (2019). Directional statistics-based deep metric learning for image classification and retrieval. Pattern Recognit.93, 113123. 10.1016/j.patcog.2019.03.008

  • 58

    ZhuL.LiaoB.ZhangQ.WangX.LiuW.WangX. (2024). Vision mamba: efficient visual representation learning with bidirectional state space model, arXiv preprint arXiv:2401.09417. 10.48550/arXiv.2401.09417

  • 59

    ZophB.VasudevanV.ShlensJ.LeQ. V. (2018). “Learning transferable architectures for scalable image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 86978710. 10.1109/CVPR.2018.00907

Summary

Keywords

big data, geological energy resources, intelligent simulation, risk early warning, underground spaces

Citation

Guo Z (2026) A distribution-aware Riemannian transformer for intelligent simulation and early warning of risks in geological energy and underground spaces. Front. Earth Sci. 14:1822287. doi: 10.3389/feart.2026.1822287

Received

04 March 2026

Revised

26 June 2026

Accepted

29 June 2026

Published

30 July 2026

Volume

14 - 2026

Edited by

Shruti Kanga, Central University of Punjab, India

Reviewed by

Awad Sohaib R., Ninevah University, Iraq

Sartsin Phakdimek, Suranaree University of Technology, Thailand

Updates

Copyright

*Correspondence: Zijun Guo,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics