ORIGINAL RESEARCH article
Front. Genet.
Sec. Epigenomics and Epigenetics
FusDRM-m5C: a hybrid model for accurate prediction of 5-Methylcytosine modification sites based on feature fusion and attention mechanism
Provisionally accepted- School of Information Engineering, Jingdezhen Ceramic University, Jingdezhen, China
Select one of your emails
You have multiple emails registered with Frontiers:
Notify me on publication
Please enter your email address:
If you already have an account, please login
You don't have a Frontiers account ? You can register here
Introduction: The precise identification of 5-methylcytosine (m5C), an epitranscriptomic modification fundamental to RNA function, is crucial yet proves difficult to achieve experimentally. Consequently, computational prediction offers a promising avenue; however, refining its predictive accuracy and ensuring its robustness remain ongoing objectives. To address these limitations, this study introduces a deep learning framework designed for highly accurate m5C site prediction from RNA sequences. Methods: We propose FusDRM-m5C, a deep learning framework featuring a multi-branch architecture designed to process three distinct feature types: one-hot vector representation (one-hot), Z-curve-based geometrical features (Z-curve), and local RNA secondary structure (RSS). Each feature type is processed by a separate, parallel branch. Within each branch, a Dilated Convolutional Neural Network (DCNN) captures multi-scale patterns, followed by a Multi-Head Self-Attention (MHSA) mechanism with residual connections to weigh context-dependent information. For feature fusion, the high-level representations from the three branches are then integrated via concatenation. This fused feature vector is subsequently fed into a final fully connected network, which generates the prediction probability for precise m5C site identification. Results: The performance of FusDRM-m5C was rigorously evaluated using both 5-fold cross-validation (CV) and independent dataset testing. On the 5-fold CV benchmark dataset, the model achieved high predictive accuracy, reflected by a Sensitivity (Sn) reaching 0.995, Specificity (Sp) of 0.971, Accuracy (ACC) at 0.983, Matthews correlation coefficient (MCC) measuring 0.966, and an Area Under the Receiver Operating Characteristic Curve (AUC) of 0.997. Crucially, when assessed on an independent test dataset, the model maintained strong generalization capability, attaining an Sn of 0.900, Sp of 0.965, Acc of 0.933, MCC of 0.867, and an AUC of 0.986. Furthermore, we assessed the cross-species prediction performance of FusDRM-m5C. The results demonstrated that the model consistently maintained high accuracy and robustness across datasets from multiple species, outperforming several existing state-of-the-art methods. Discussion: The proposed FusDRM-m5C model demonstrates highly accurate and robust prediction of m5C sites, comparing favorably with existing methods. Its architecture effectively integrates diverse biological features through distinct processing pathways fused via attention, offering a powerful tool for m5C identification. The source code and experimental data are available at https://github.com/hhui0/FusDRM-m5C-code.
Keywords: m5C site identification, Multi-feature fusion, deep learning, Dilated ConvolutionalNeural Networks, multi-head self-attention
Received: 10 Jun 2025; Accepted: 03 Nov 2025.
Copyright: © 2025 Huang, Jia and Zhou. This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
* Correspondence: Jianhua Jia
Disclaimer: All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.
