ORIGINAL RESEARCH article

Front. Microbiol., 25 September 2025

Sec. Microorganisms in Vertebrate Digestive Systems

Volume 16 - 2025 | https://doi.org/10.3389/fmicb.2025.1660775

Gut microbiota analysis reveals microbial signature for multi-autoimmune diseases based on machine learning model

  • TA

    Tianfeng An 1

  • SZ

    Shuya Zhang 2

  • JL

    Jinjin Li 3

  • HW

    Hui Wang 1,4,5

  • LC

    Li Chen 1,4,5,6

  • YS

    Yiran Shi 1

  • JW

    Jingyi Wang 7

  • SH

    Sirui Han 7

  • RW

    Ruoxi Wang 1

  • LW

    Linyuan Wang 1

  • ZH

    Zijing Huan 1

  • RY

    Ruiqi Yang 1

  • DH

    Desong Hao 1

  • YL

    Yanfang Liu 1

  • XL

    Xuehua Liu 3† *

  • CY

    Chao Yuan 1,4,5,6† *

  • 1. The School of Public Health, Tianjin Medical University, Tianjin, China

  • 2. The School of Medicine, Nankai University, Tianjin, China

  • 3. Department of General Practice, Tianjin Union Medical Center, The First Affiliated Hospital of Nankai University, Tianjin, China

  • 4. Tianjin Key Laboratory of Environment, Nutrition and Public Health, Tianjin Medical University, Tianjin, China

  • 5. Center for International Collaborative Research on Environment, Nutrition and Public Health, School of Public Health, Tianjin Medical University, Tianjin, China

  • 6. Department of Toxicology and Sanitary Chemistry, School of Public Health, Tianjin Medical University, Tianjin, China

  • 7. The School of Clinical Medical, Tianjin Medical University, Tianjin, China

Abstract

Introduction:

Human microbiota is a major factor contributing to the immune system, offering an opportunity to develop non-invasive methods for disease diagnosis. In some research on Autoimmune Diseases (AIDs), gut microbiota variation has been observed. However, there remains a paucity of research that explores the potential of gut microbiota as a microbial signature for the classification and diagnosis of multi-AIDs.

Methods:

In this study, we analyzed 1,954 gut microbiota sequencing datasets from public databases collected from 1,043 patients with 10 AIDs to identify common or unique microbial signatures for AIDs through differential abundance testing and machine learning techniques. We evaluated five popular algorithms: Random Forest (RF), Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Multilayer Perceptron (MLP), and eXtreme Gradient Boosting (XGBoost) models. Five-fold cross-validation and grid search were used to select the model parameters.

Results:

After comparing the performance of five models, the XGBoost model showed superior performance and achieved an area under the receiver operating characteristic curve (AUROC) ranging from 0.75 to 0.99 when predicting different diseases in the test set. At a specificity of 0.7 to 0.96, the sensitivity ranged from 0.66 to 1. By correlating the top 77 microbiota genera with the disease phenotypes, 126 significant associations were identified [false discovery rate (FDR) < 0.05]. We improved the detection accuracy and disease specificity for AIDs and revealed microbiota features specific to 10 different AIDs. Moreover, we found changing trends in shared microbiota features across some AID phenotypes, such as Crohn's Disease (CD) and Ulcerative Colitis (UC). At the same time, opposite changing trends were observed in the shared microbial signatures, such as Psoriasis and Myasthenia Gravis (MG). These results suggest that specific gut microbiota genera may affect the host immunity and induce different AID phenotypes.

Discussion:

This research holds potential for clinical application in the auxiliary diagnostic evaluation and monitoring of treatment responses. Simultaneously, it provides important clues for research on the characteristics of the intestinal immune microenvironment for different AIDs.

Introduction

The gut microbiota represents a highly intricate community of microorganisms, encompassing bacteria, archaea, and eukaryotes, that inhabit the intestine. It is estimated that in approximately 100 trillion cells (), the symbiotic microbes within the human gut exceed the number of host cells by at least an order of magnitude and possess a substantially larger repertoire of unique genes when compared to the host genome. Overall, the intestinal microbiota is composed of approximately 500–1,000 species, which belong to only a limited number of known bacterial phyla (Zoetendal et al., 2008). Emerging research suggests that the gut microbiota demonstrates vast enzymatic potential and plays a pivotal role in modulating diverse aspects of host physiology, such as pathogen resistance, host immunity, and metabolic processes (; ; ).

The gut microbiota plays a crucial role in maintaining a delicate equilibrium between host defense and immune tolerance (Yang et al., 2021). The microenvironment within the gut is shaped by complex and intricate interactions between the gut microbiota and the local innate immune system (; ). Optimally, the interaction between the immune system and microbiota functions in harmony, integrating both the innate and adaptive components of immunity to select, calibrate, and terminate immune responses, thus preserving homeostasis. Nevertheless, this immune balance between the gut flora and host is not invariably stable. A variety of pathologies affecting humans, including allergies, autoimmune disorders, and inflammatory conditions, stem from the inability to regulate misdirected immune responses against self-antigens, microbiota-derived antigens, or environmental antigens (Zoetendal et al., 2008; ; ; ).

Autoimmune diseases (AIDs) occur when an individual's immune system mistakenly attacks its own tissues, with an estimated global incidence of approximately 3–5% (; ). Human microbiota is thought to be a crucial factor in the development of autoimmunity, as alterations in the microbial composition can result in the breakdown of immune tolerance (; ). Systematic analysis of the human gut microbiota holds promise for developing non-invasive diagnostic methods for major AIDs. With the emergence of next-generation sequencing (NGS), novel strategies have been developed to investigate the association between gut microbiota dysregulation and AIDs. These strategies entail bioinformatic analysis to characterize the microbial compositions of samples, including the identification of microbial taxa and their relative abundance. Moreover, in case–control studies, researchers have attempted to identify differentially abundant microbial taxa as potential disease biomarkers (; ; ). This approach has been applied to a variety of AIDs, such as systemic lupus erythematosus (SLE), rheumatoid arthritis (RA), inflammatory bowel diseases (IBD), and systemic sclerosis (SS) (; ; ; ; ; ; ; ).

However, the existence of shared microbial signatures across diverse diseases and overlapping gut microbiota signatures among most health states poses challenges for accurate diagnosis when using single-disease models, which can result in misclassification (). To address this, multi-class diagnostic models have been developed to predict disease-specific signatures across the microbiota, enabling more accurate diagnostic purposes (; ; ). Machine learning (ML) classifiers are often employed for disease diagnosis, either using gut microbiota data alone or in conjunction with clinically relevant features, to differentiate patients from healthy controls (; ; ). ML-based gut microbiota classifiers have been developed for a variety of diseases, including inflammatory bowel disease (IBD), liver cirrhosis (LC), autism spectrum disorder (ASD), Alzheimer's disease (AD), and numerous others (; ; ; ; ; ; ; ).

In the present study, we conducted a comprehensive meta-analysis of multiple AIDs. A total of 1,954 samples were used in this study. These diseases spanned 10 major disease categories. To comprehensively characterize the gut microbiota in relation to 10 AIDs, namely rheumatoid arthritis (RA), ankylosing spondylitis (SpA), multiple sclerosis (MS), psoriasis, Crohn's disease (CD), ulcerative colitis (UC), celiac disease (CeD), myasthenia gravis (MG), systemic lupus erythematosus (SLE), and type 1 diabetes (T1D), we utilized gut microbiota sequencing data to evaluate the abundance of taxonomic units. Furthermore, we developed a machine learning multi-classification model for the diagnosis of multi-AIDs and identification of microbial signatures that are either common across or specific to these 10 AIDs.

Materials and methods

The main framework for dataset partitioning, model training, and validation of this research was shown in Figure 1A.

Figure 1

Data collection

A comprehensive list of human autoimmune disease-related case–control studies on gut microbiota was performed in public databases, including NCBI BioProject (https://www.ncbi.nlm.nih.gov/bioproject) and GMrepo (a curated database of human gut metagenomes; https://gmrepo.humangut.info) (; Wu et al., 2020). A total of 1,954 gut microbiota sequences on 10 different autoimmune diseases (Supplementary Table S1, run-level data including 1,043 cases and 911 controls) were collected. AIDs included RA, SpA, MS, Psoriasis, CD, UC, CeD, MG, SLE, and T1D. The specific inclusion and exclusion criteria are as follows: our inclusion criteria included: (1) Study on 16S rRNA amplicon sequencing of gut microbiota with complete disease phenotype metadata. (2) Case–control studies with at least nine valid samples in each case and control group. (3) No antibiotics or probiotic supplements were administered in the past 3 months. The exclusion criteria were as follows: (1) post-treatment follow-up after medication. We divided these 10 diseases into 7 categories, including 3 digestive system diseases, 2 AIDs of the nervous system diseases, 2 musculoskeletal diseases, 1 endocrine system disease, 1 connective tissue disease, and 1 skin disease, according to the NCBI Medical Subject Headings (MeSH, https://meshb.nlm.nih.gov/) database and Human Disease Ontology (DO) database ().

Sequencing data extraction and microbiota profiling

We extracted all SRA_ID (listed in Supplementary Table S1) from the NCBI SRA database (https://www.ncbi.nlm.nih.gov/sra) () or European Nucleotide Archive (ENA) () and obtained related information on samples from the NCBI BioSample database (https://www.ncbi.nlm.nih.gov/biosample). Raw sequencing data (FATSQ files) were then downloaded using the SRA toolkit, and metadata, including age, gender, and country, were collected. Trimmomatic () was used to trim the reads and to remove sequencing vectors and low-quality bases. Reads shorter than 50 base pairs were discarded after trimming. To preprocess the sequences, we used QIIME (2023.5) () to demultiplex and quality-filter the data. Representative sequences and their abundances were extracted using the feature table () to generate tables containing amplicon sequence variants (ASVs). Taxonomic assignment of the individual dataset was classified against the Greengenes database (version 13.8123) (). Genus-level relative abundance results were retained for subsequent analyses. Subsequently, samples with only two or fewer taxa were excluded from subsequent analyses. Additionally, to minimize the noise caused by low-abundance taxa, we filtered out those with a relative abundance of < 0.001 across all samples.

Microbiota analysis

All statistical analyses were performed using R, version 4.2.2. The ggpubr package (https://github.com/kassambara/ggpubr) was used to perform non-parametric statistical testing between groups and to account for multiple hypothesis testing corrections when necessary. Significant differences in alpha diversity (Shannon index and richness) were determined using the non-parametric Kruskal–Wallis test. Differences in beta diversity (Bray–Curtis distance matrix calculated using the relative abundances of microbial genera) were determined by PERMANOVA using distance matrices (adonis) in the adonis function of the vegan R package with 999 permutations. Principal coordinate analysis (PCoA) based on beta-diversity was used to visualize the clustering of samples based on their genus-level compositional profiles. To adjust these findings for other factors that may affect microbiota, we used microbiota multivariable associations with linear models (MaAsLin) to identify compositional differences while adjusting for age, sex, and country. To account for the sequencing batch effects of all samples treated at different periods, we used the adjust_batch function implemented in the “MMUPHin” R package using project ID as the controlling factor before model development. Detailed information on the effectiveness of the batch effect removal in this study is shown in Supplementary Figure S1.

Classification model for multiple-autoimmune-diseases

Binary sub-cohorts were composed of one AID phenotype and its corresponding healthy control, resulting in a total of 10 binary subgroups. The random forest (RF) model was chosen as the binary classifier because its classification performance has been shown to outperform other methods for microbiota data (). The RF model was first trained on a randomly selected training set (5-fold stratified cross-validation) and then applied to the withheld test set to assess the final performance. This process was repeated 20 times to obtain a distribution of RF prediction evaluations for the test set, and the mean AUROC value was calculated.

Multi-class models can be implemented in Python 3.8 using standard libraries that are publicly available, including pandas (2.0.3), numpy (1.22.4), scikit-learn (1.3.1), and matplotlib (3.7.3). For each phenotype, the samples were randomly divided into training (70%, n = 1,368) and test (30%, n = 586) sets. Random forest (RF), support vector machine (SVM), K-nearest neighbors (KNN), multilayer perceptron (MLP), and eXtreme Gradient Boosting (XGBoost) were used as classifier models to diagnose different phenotypes based on the taxonomic profiles at the genus level of the gut microbiota. We employed a 5-fold cross-validation and grid search to select the optimal model parameters for the RF, SVM, KNN, and XGBoost models. For the MLP models, we implemented the default settings provided by Scikit-learn. Finally, we evaluated the performance of the five models on the test dataset as the final performance for predicting different diseases. The highly ranked and frequently selected microbiota features were considered predictive signatures for further interpretation.

Model validation and performance evaluation

The area under the receiver operating characteristic curve (AUROC) is a widely used metric for evaluating the performance of classification models and provides a comprehensive assessment of the model's sensitivity and specificity trade-off at different thresholds. The range of AUROC values is typically explained from 0.5 to 1, with a higher value indicating a better ability to distinguish between different classes of samples. The area under the precision-recall curve (AUPRC), as a complementary assessment, considers the trade-offs between precision (or positive predictive value) and recall (or sensitivity) and is more robust for imbalanced datasets. The AUPRC ranges from 0 to 1, with a value of 0 signifying no positive examples identified and a value of 1 indicating perfect identification of all positive examples. In addition, for a more comprehensive evaluation of the model's performance, we employed the F1 score. It is a metric that considers the precision and recall of a model. The F1 score ranges from 0 to 1, with a higher value indicating better overall performance of the model.

We employed the bootstrap method to address the data imbalance and obtain more robust performance evaluations for each model. The bootstrap method involves iteratively resampling the training data, training a new model, and evaluating the model 1,000 times (; ). The performance of the model was calculated as the average performance of the individual models developed using the bootstrap method. Bootstrap methods can considerably reduce overfitting in the developed models. To identify the most discriminative features among many bacterial genus features, minimize model complexity, and enhance computational efficiency and interpretability, we generated a learning curve relating the number of features to the model performance.

Sensitivity analysis

Considering the potential influence of the three factors, “gender,” “age,” and “geographical location,” on the gut microbiota, we conducted the sensitivity analysis. For the factor “geographical location,” >75% of the samples in this study were collected from the United States (746) and China (724). We evaluated our model using country-based stratification. For factors “gender” and “age,” the Kruskal–Wallis test was conducted to verify the significance of the differences in the abundance of the corresponding genus in the gut microbiota among different age groups and different genders regarding the AID phenotype. Age groups: Juvenile (3–20), Young Adult (21–45), Older Adult (46–65), and Elderly (65+); Gender groups: Female and Male (Supplementary Table S1, “Sensitivity analysis”).

Results

Summary of available data

A total of 1,954 gut microbiota sequencing data (1,043 AIDs cases and 911 non-disease controls; sequences based on Illumina platforms) were collected from the NCBI database based on 19 case–control studies. These data could contain up to 127 cases and 247 controls; however, most studies were conducted on a limited number of samples, with median sizes of 55 cases and 48 controls, respectively. The fecal samples for the gut microbiota sequencing data were mainly collected from five continents and 13 countries, most of which were from the United States of America (38.18%), China (32.44%), and Canada (11.36%) (related information was shown in Supplementary Table S1, “Data availability”).

Gut microbiota features across different phenotypes

Ecological indices may not be robust indicators for distinguishing disease from healthy status, which results in changes in the structure of the gut microbiota. First, we aimed to determine the differences in the composition of the gut microbiota among the different AIDs. Compared to healthy controls, significant differences in bacterial diversity (Shannon) and richness (number of genera) were observed in AIDs, except for RA. The Shannon index of the digestive system AIDs decreased. Moreover, we found that both indices (Shannon and richness) varied across phenotypes (Figure 1B). The results of microbial diversity among different phenotypes are consistent with a recent study on multi-class disease diagnosis based on gut microbiota and meta-analysis (; ). Beta diversity based on Bray–Curtis dissimilarity showed significant differences in gut microbiota composition among individuals with different phenotypes (R = 0.396; F = 36.057; p < 0.001) (Figure 1C). We then used the linear model of MaALin2 after adjusting for age, sex, country, and technical confounders to explore the associations of microbial composition at the genus level with disease phenotypes. We found 192 significant associations between the 11 phenotypes and 62 bacterial taxa at the genus level (FDR < 0.05). Among the 62 genera, > 67.7% were significantly associated with two or more diseases. The genera Haemophilus, Veillonella, Eggerthella, and Rothia were positively associated with AIDs in our results, whereas the opposite was observed with the genera Paraprevotella, SMB53, and Gemmiger (Figure 1D, Supplementary Table S1 “Gut microbiota data”).

Development of a gut microbiota-based diagnosis model

The binary classifier of the RF model based on the microbiota could significantly distinguish between healthy and most AIDs (Figure 2A), which indicated that AIDs had different degrees of intestinal flora disturbance compared with the control. To further investigate the discriminatory ability of the gut microbiota in various AIDs and to distinguish between AIDs and healthy controls, we constructed multi-class classifiers.

Figure 2

To select the best multi-class machine learning algorithm, we evaluated five popular algorithms: RF, SVM, KNN, MLP, and XGBoost models. Five models achieved a mean AUROC of all phenotypes of 0.72–0.89 (Figure 2B), suggesting that muti-class disease classification based on the gut microbiota was feasible. Amongst them, the XGBoost multi-class model had an optimal performance and achieved a mean AUROC of 0.89 [interquartile range, IQR (0.87–0.90)], a mean AUPRC of 0.48 [interquartile range, IQR (0.44–0.51)], and a mean F1-score of 0.538 [interquartile range, IQR (0.51–0.57)] for different disease phenotypes in the test set (Figure 2C). Therefore, the XGBoost multi-class classifier was used for further analysis.

The XGBoost model is constructed using a complete training set. The importance rankings of all features at the genus level were obtained by leveraging this model. Subsequently, the features were incorporated in descending order of importance. For each step of adding a feature, the AUROC value of the model was computed, thereby generating a learning curve that depicted the relationship between the number of features and model performance (Figure 2D). The results indicate that the model performance reached a plateau when 77 features were employed. At this stage, a further increase in the number of features did not lead to a significant improvement in performance. Consequently, the first 77 features were selected as the input variables for the final model. The AUROC values for most phenotypes exceed 0.9 (Figure 3A). The macro-AUROC value was 0.9 [IQR (0.88–0.91)], indicating superior performance compared to the binary classifier we trained. This classifier proved to be valuable for effectively distinguishing AIDs based on features derived from the gut microbiota. At the optimal Youden's index threshold, the sensitivities of our XGBoost multi-class model range from 0.66 to 1 at specificities of 0.70 to 0.95 for different diseases with accuracy from 0.74 to 0.94, highlighting good diagnostic performance (Figure 3B). For example, our model achieved an AUROC of 0.95, CD with a sensitivity of 0.94, and a specificity of 0.89 (Figures 3A, B). To better characterize the XGBoost model, we compared its performance on datasets with different split ratios and obtained similar results, which indicated that the model exhibited high stability and good predictive capability without the risk of overfitting (Figure 3C).

Figure 3

Associations between microbiota features and phenotypes

Next, we correlated the top 77 bacterial genera that contributed to the model with the different disease phenotypes. Among the 77 genera, 126 significant associations were found between 42 genera and the different disease phenotypes. These 42 genera belonged to the phyla Firmicutes (28 genera), Actinobacteria (6 genera), Proteobacteria (5 genera), Fusobacteria (1 genus), and Bacteroidetes (3 genera) (Figure 3D). Among the selected bacterial genera, significant associations were found between the phenotypes and several genera, except for T1D (genera only). From the perspective of AID phenotypes, CD (23 genera), RA (21 genera), and psoriasis (20 genera) were the three phenotypes with the highest number of related genera (FDR < 0.05). However, in SpA and CeD, fewer related genera were found (n < 5). From the perspective of the gut microbiota, the genera Actinobacteria and Ruminococcaceae II correlated with the largest number of AID phenotypes (six AID phenotypes). The genera Shigella, Clostridium, and Eggerthella also correlated with many AID phenotypes (five AID phenotypes). These genera could serve as shared microbiota features, which may suggest an association with AID phenotypes except for T1D. The genus Dorea, Lachnobacterium, WAL_1855D, Bulleidia, Pseudomonas, and the special genus Prevotella are only associated with one AID phenotype. This may imply that specific changes in the gut microbiota are related to the corresponding AID phenotype. The genera Clostridium, Eggerthella, Haemophilus, Fusobacterium, Subdoligranulum, and Rothia positively correlated with several AID phenotypes. However, Paraprevotella, SMB53, Clostridium II, Gemmiger, and Slackia were negatively correlated with several AID phenotypes. Only Fusobacterium, which belongs to Fusobacteria, was positively correlated with two inflammatory bowel diseases (CD and UC), indicating a potential microbiota feature of inflammatory bowel diseases.

Interestingly, we noted a higher degree of similarity in microbial alterations between diseases within the same system, such as inflammatory bowel diseases (CD and UC). Among these two phenotypes, 11 genera (approximately 50%) were shared between these two phenotypes, exhibiting similar trends in microbial changes. CeD, also an autoimmune disease of the digestive system, presents a completely distinct profile of related genera. Analogous results have also been reported in Psoriasis and SLE. Although shared microbiota features were relatively scarce, eight genera with similar trends in microbial changes were identified. In autoimmune diseases of the nervous system (MS and MG) and musculoskeletal system (RA and SpA), such similarities in microbial changes were not observed. On the other hand, in some AID phenotypes, alterations in the microbiota showed an opposite trend compared to the healthy group. For example, in Psoriasis and MG, ten microbiota features were shared between the two AID phenotypes, including the genera Actinobacteria, Butyricicoccus, Clostridium IV, Blautia, Lactonifactor, Shigella, Anaerostipes, Clostridium I, Parabacteroides, and Ruminococcus II. However, opposite trends were observed. Similar findings were obtained in a comparison of psoriasis and RA. These results suggest that, in different AIDs, the gut microbiota microenvironment may possess completely opposing characteristics.

Sensitivity analysis

The data collected for this study were sourced from multiple countries. More than 70% of the study participants were from the United States and China. The model's performance was evaluated using country-based stratification, revealing consistent results (Figure 4). Moreover, the model attained a mean Area Under the Receiver Operating Characteristic Curve (AUROC) of 0.90 (Interquartile Range, IQR: 0.88–0.93) in the United States and 0.91 (0.88–0.93) in China, respectively. Given the potential influence of “gender” and “age” on the gut microbiota, we performed a stratified analysis based on age and gender (Supplementary Table S1, “Sensitivity analyses”). The sensitivity analysis results indicated that “Geographic location,” “Age,” and “Sex” are not the primary factors affecting classification outcomes.

Figure 4

Discussion

The human gut microbiota has assumed growing significance as a biomarker for non-invasive disease screening and as a target for disease intervention owing to its profound association with human diseases. In this study, we comprehensively aggregated publicly available datasets of gut microbiota sequencing. Moreover, we integrated microbiota features among diverse AIDs and utilized advanced and reproducible machine learning approaches that are highly relevant to clinical practice. Overall, this study demonstrated the feasibility of using a multi-classifier based on gut microbiota for the identification of AIDs. We propose that this multi-class model shows great potential for clinical applications in auxiliary diagnosis evaluation and assessment of the efficacy of interventions. Simultaneously, it offers crucial clues for the investigation of the characteristics of the intestinal immune microenvironment in different AIDs that primarily affect various body sites.

Our analysis of the gut microbiota revealed microbiota features associated with the 10 AIDs. Most of these findings are consistent with previous research on the correlation between the gut microbiota and AIDs. For example, Actinomyces spp., which act as both pathogens and constituents of human microbiota, have been reported to be enriched in patients with IBD or colorectal cancer (Yachida et al., 2019; ). Bifidobacterium spp. are important probiotics, particularly during infancy. In the infant gut microbiota, a relationship between Bifidobacterium spp. and allergies has been identified, which plays a vital role in immune modulation and tolerance (; ; ). Adlercreutzia spp. and Clostridium spp. were decreased in the gut microbiota of patients with CD (). Haemophilus spp., a type of gut microbiota, has been reported to be associated with several immune-related disorders, including RA, IBD, and Hashimoto's thyroiditis (; Zhang et al., 2015; ). Similarly, in our study, we also detected significant enrichment of this bacterial genus in patients with RA and IBD. Meanwhile, in studies exploring the gut microbiota of patients with type 2 diabetes (T2D), Parkinson's disease, and Alzheimer's disease, differences in the genus Haemophilus were observed between the disease groups and healthy control groups (; ). Eggerthella spp. and Prevotella spp., as typical gut microbiota signatures, were reported to be enriched in SLE, which is consistent with our results (; Yao et al., 2023). Genus Eggerthella levels have been implicated in inflammatory diseases, especially human gut Eggerthella lenta in AIDs and other conditions (; Xiang et al., 2021; ). Veillonella spp. are closely related to the genus Clostridiales, which are recognized as probiotic organisms in the host (). Other studies have indicated that Veillonella spp. act as inflammophilic pathobionts, thriving in an inflammatory milieu, and possess the inherent ability to stimulate IL-6-mediated inflammation (). Notably, in a study of the gut microbiota in autoimmune hepatitis (AIH), the genus Veillonella was identified as the key genus strongly associated with the disease (; Yuming et al., 2024). The genus Fusobacterium was significantly increased in IBD in our study, and previous research has found a similar association (; ; ; ; ). In particular, Fusobacterium nucleatum exhibits a wide range of characteristics under certain conditions. It can adhere to a large number of phylogenetically unrelated bacterial species, potentially leading to the translocation of non-invasive, yet pro-inflammatory species across the compromised intestinal epithelium, thereby exacerbating the disease state (; ). Given that the AIDs included in this study preferentially targeted distinct body organs, it is unsurprising that we detected differences, as prior studies have reported analogous conclusions (). In our study, the number of microbiota features corresponding to each AID phenotype was lower than the number of significant differential microbiota abundances found in existing studies based on basic case–control studies. This is attributable to the fact that our model was constructed using data from 10 AIDs. In addition to considering the differences between AIDs and controls, it also accounts for potential disparities among the various AID phenotypes. Consequently, the selected microbiota features are relatively fewer. However, some of our findings deviated from those of previous studies. For example, previous studies have indicated that Prevotella spp. are more abundant in the early/preclinical stages of RA, whereas Bifidobacterium spp. are less abundant. Nevertheless, our study did not yield similar results (). Our findings also identified several bacterial genera that were scarcely observed in previous investigations of the 10 AIDs. Rothia spp. have been mentioned in research on treatable periodontitis, endocarditis, and joint infections (; ; ). Our study revealed a potential association between the genus Rothia and SLE, RA, and IBD, which has not been previously reported. Paraprevotella spp., encompassing numerous opportunistic pathogens, has only been reported in the fecal samples of patients with rare AIDs, such as Behcet's disease (BD), Vogt-Koyanagi-Harada (VKH) disease, and the dextran sulfate sodium-induced IBD model (Ye et al., 2020, 2018; ). In our study, there was a significant reduction in Paraprevotella in MS, CD, UC, and CeD. In the colorectal cancer mouse model, the Paraprevotella spp.-derived metabolite agmatine triggered inflammation to promote colorectal tumorigenesis through the Wnt signaling pathway (). This might be the key mechanism by which Paraprevotella spp. regulates the gut immune microenvironment in various AIDs. Although disease phenotypes were not observed in this study, the genus Slackia was found to be more abundant in patients with APECED-associated hepatitis (APAH) (). The discovery of these bacterial genera suggests that multi-classifiers based on deep machine learning are more conducive to uncovering gut microbiota features that have not been easily discerned in previous studies across a broader range of data.

As described in the “Results” section, we observed a higher degree of similarity in microbial alterations among the two inflammatory bowel diseases (CD and UC). Among these two phenotypes, 11 shared genera exhibited similar trends in terms of microbial changes. Analogous findings were also noted in Psoriasis and SLE, where eight shared genera with similar trends in microbial changes were identified. In contrast, in the nervous system, AIDs (MS and MG) and musculoskeletal system (RA and SpA) AIDs, such similarities in microbial changes were not detected. IBD is a consequence of the interaction between the host and microorganisms, encompassing intestinal microbial factors, abnormal immune responses, and a damaged intestinal mucosal barrier. CD and UC are two subtypes of IBD that may have similar intestinal immune microenvironments, leading to numerous shared microbiota features with comparable trends of microbial changes (; ). Psoriasis and SLE are chronic autoimmune diseases that affect multiple organs. Although their specific pathogenic mechanisms differ, they all involve abnormal activation of immune cells, abnormal secretion of cytokines, and T cell-mediated inflammation (; ). In our results, eight genera, Eggerthella, Subdoligranulum, Ruminococcus, Lactonifactor, Anaerotruncus, Clostridium, Parabacteroides, and Shigella, shared microbiota features with similar trends. Eggerthella spp. can induce intestinal Th17 activation by lifting inhibition of the Th17 transcription factor Rorγt through cell- and antigen-independent mechanisms (). Subdoligranulum spp. are arthritogenic strains that trigger RA and can stimulate Th17 cell expansion in mice (Zhu et al., 2020). In colorectal cancer, Ruminococcus spp. can maintain the immune surveillance function of CD8+ T cells through its metabolic characteristics (). The remaining three genera were also associated with other AIDs, including MS, RA, and IBD (; ; ; Yousefi et al., 2023; ).

On the other hand, in some AID phenotypes, alterations in the microbiota exhibited an opposing trend compared to the healthy control group. For example, in Psoriasis and MG, 10 microbiota features were shared between these two AID phenotypes, including the genera Actinobacteria, Butyricicoccus, Clostridium IV, Shigella, Blautia, Lactonifactor, Anaerostipes, Clostridium I, Parabacteroides, and Ruminococcus II; however, opposite trends of change were found. Similar results were also found in Psoriasis and RA (opposite-trend-genera: Actinobacteria, Clostridium IV, Blautia, Clostridium I, Ruminococcus I, Ruminococcus II, Parabacteroides, Bilophila Lactonifactor, Megamonas, Anaerostipe, Holdemania, and Shigella). Th17 cells play a crucial role in the pathogenesis of AIDs. Although specific pathological regions vary among different AID phenotypes, Psoriasis, RA, and MG involve an imbalance in Th17/Treg cells (; Zhang et al., 2024; ). Almost all the above-mentioned genera with opposite trends have been reported to be associated with Th17 cells or T-cell-related inflammatory responses, except for the genera Holdemania and Bilophila (; ; ; ; ; Yu et al., 2025; ; Zou et al., 2024). Numerous investigations have demonstrated that the gut microbiota can exert an impact on host immunity via its metabolites. This process, in turn, affects AID phenotypes, and this mechanism represents a key strategy for treating AIDs (Yang and Cong, 2021). Our findings indicate that, within distinct AIDs, the gut microbiota microenvironment may exhibit completely opposing characteristics. The variations in AID phenotypes may be attributed to the influence of specific gut microbiota patterns on the host immune response process. Combined with the role of genetic factors, this leads to an imbalance of Th17/Treg cells in different regions of the body, ultimately giving rise to the emergence of corresponding pathological phenotypes.

Strengths and limitations

The strengths of this study include the use of gut microbiota data covering a variety of disease phenotypes (10 AIDs), including gut microbiota sequencing data from almost 2,000 participants. However, this study has some limitations that should be acknowledged. First, we aimed to include gut microbiota sequencing data covering the widest possible range of AIDs and the corresponding healthy controls. However, microbiota sequencing data in public databases often lack relevant information, such as host comorbidities, diet, BMI, and treatment/medication conditions. In our study, only sex (70.5%), age (86.9%), and geographical location (100%) were comprehensively available. Previous research has shown that age and geographical location influence gut microbiota composition, and in some diseases, such as systemic lupus erythematosus (SLE), sex also plays a role (; ; ). Therefore, we conducted sensitivity analyses based on these three factors (Figure 4 and Supplementary Table S1, “Sensitivity analyses”). Most genera showed significant differences; however, in the ≥65 age group, some genera did not, likely due to the limited sample size rather than age-related effects. Moreover, because of the limitations of publicly available sequencing data (16S limitations), our analyses were restricted to the genus level, and species-level analyses could not be performed. Second, all datasets were retrospectively obtained, and the classifier was not validated in an external prospective cohort. The complex phenotypes of AIDs, prolonged sample collection periods, and limited availability of prospective data in public databases preclude such validation. Future studies should involve collaboration with relevant clinical teams to prospectively collect fecal samples, utilize gut metagenomic sequencing to obtain deep-level microbiota information, assess the classifier's predictive value for disease progression or prognosis, and enhance its clinical applicability. Finally, the data were cross-sectional, limiting our ability to determine the temporal sequence and causality between abnormal gut microbiota and the onset of AIDs.

Conclusion

In conclusion, this study offers a comprehensive analysis of the composition of stool microbiota in AIDs. Our findings show that the composition of gut microbiota changes in rheumatoid arthritis (RA), spondyloarthritis (SpA), Crohn's disease (CD), ulcerative colitis (UC), celiac disease (CeD), multiple sclerosis (MS), myasthenia gravis (MG), systemic lupus erythematosus (SLE), and psoriasis. These changes are notably associated with varying degrees of gut dysbiosis. Moreover, through differential abundance testing and machine learning techniques, we identified several microbial signatures that exhibit consistently higher or lower abundances in AIDs patients than in healthy controls. Subsequent research is imperative to delve into the specific roles and functions of this genus within the host. This is crucial for establishing causal associations in disease pathogenesis and for exploring their potential as targets for future therapeutic interventions.

Statements

Data availability statement

The data that supports the findings of this study are available in NCBI SRA database (https://www.ncbi.nlm.nih.gov/sra); accession numbers listed in Supplementary Table S1.

Author contributions

TA: Writing – review & editing, Investigation, Resources, Software, Supervision, Writing – original draft, Data curation, Validation, Formal analysis, Visualization, Methodology. SZ: Validation, Methodology, Data curation, Investigation, Conceptualization, Resources, Writing – original draft. JL: Data curation, Investigation, Resources, Funding acquisition, Writing – review & editing. HW: Resources, Data curation, Writing – review & editing, Conceptualization. LC: Resources, Project administration, Data curation, Writing – review & editing. YS: Data curation, Resources, Writing – review & editing. JW: Writing – review & editing, Resources, Data curation. SH: Writing – review & editing, Resources, Data curation. RW: Data curation, Resources, Writing – review & editing. LW: Writing – review & editing, Resources, Data curation. ZH: Resources, Data curation, Writing – review & editing. RY: Data curation, Resources, Writing – review & editing. DH: Data curation, Writing – review & editing, Resources. YL: Data curation, Resources, Writing – review & editing. XL: Supervision, Data curation, Resources, Funding acquisition, Project administration, Writing – review & editing. CY: Formal analysis, Resources, Writing – review & editing, Funding acquisition, Visualization, Project administration, Writing – original draft, Supervision, Software, Methodology, Data curation, Conceptualization, Investigation, Validation.

Funding

The author(s) declare that financial support was received for the research and/or publication of this article. This work was funded by the National Natural Science Foundation of China (NSFC) Program (Grant No. 82273688) and the Tianjin Municipal Education Commission (No. 2020KJ202 and No. 2022KJ202).

Conflict of interest

The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declare that no Gen AI was used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fmicb.2025.1660775/full#supplementary-material

Supplementary Figure S1

(A) Detailed information on the effectiveness of batch effect removal. Shrinkage of the batch mean parameters. X-axis: estimated batch mean parameter (Gamma); Y-axis: batch mean parameter after shrinkage (Gamma-shrunk). Shrinkage can stably adjust the batch mean parameters. (B) Original/adjusted mean abundance. X-axis: overall mean; Y-axis: mean values from the different batches. The correction process makes the originally scattered batch expressions more compact and closer to the overall average expression, thereby significantly reducing the batch effect.

Supplementary Table S1

Comprehensive information on autoimmune diseases (AIDs) was collected for this study. Details of gut microbiota sequencing data and related samples (Bioproject_ID, SRA_ID, AIDs, Age, Gender, Country) were listed in sheet 1, “Data availability.” Ten autoimmune diseases (AIDs) are listed in Table 1. Microbial abundance results of gut microbiota sequencing data and related ID were shown in sheet 2, “Gut microbiota sequencing.” Sensitivity analysis on two factors (gender and age) was shown in sheet 3, “Sensitivity analysis.” For each genus, the Kruskal–Wallis test was conducted to verify the significance of the differences in the abundance of the corresponding genus in the gut microbiota among different age groups and sexes regarding the AID phenotype.

References

Summary

Keywords

gut microbiota, autoimmune diseases, machine learning, microbial signature, auxiliary diagnostic evaluation

Citation

An T, Zhang S, Li J, Wang H, Chen L, Shi Y, Wang J, Han S, Wang R, Wang L, Huan Z, Yang R, Hao D, Liu Y, Liu X and Yuan C (2025) Gut microbiota analysis reveals microbial signature for multi-autoimmune diseases based on machine learning model. Front. Microbiol. 16:1660775. doi: 10.3389/fmicb.2025.1660775

Received

06 July 2025

Accepted

18 August 2025

Published

25 September 2025

Volume

16 - 2025

Edited by

Yanfen Cheng, Nanjing Agricultural University, China

Reviewed by

Yuqi Li, University of Illinois at Urbana-Champaign, United States

Shahzaib Khoso, University Hospital Maggiore of Carita, Italy

Updates

Copyright

*Correspondence: Chao Yuan Xuehua Liu

‡These authors have contributed equally to this work

†Present address: Chao Yuan and Xuehua Liu, Tianjin Medical University, Tianjin, China

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics