<?xml version="1.0" encoding="utf-8"?>
    <rss version="2.0">
      <channel xmlns:content="http://purl.org/rss/1.0/modules/content/">
        <title>Frontiers in Big Data | New and Recent Articles</title>
        <link>https://www.frontiersin.org/journals/big-data</link>
        <description>RSS Feed for Frontiers in Big Data | New and Recent Articles</description>
        <language>en-us</language>
        <generator>Frontiers Feed Generator,version:1</generator>
        <pubDate>2026-08-25T16:04:15.948+00:00</pubDate>
        <ttl>60</ttl>
        <item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1889044</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1889044</link>
        <title><![CDATA[A cognitively inspired feature-level fusion framework for interpretable retail sales forecasting using integration of extreme gradient boost machine, artificial neural network and attention mechanism model]]></title>
        <pubdate>2026-08-20T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Munienge Mbodila</author><author>Omobayo Ayokunle Esan</author>
        <description><![CDATA[IntroductionAccurate retail sales forecasting in modern data-intensive environments requires models that not only achieve high predictive accuracy but also scale efficiently as data volume and feature complexity increase. This study proposes a novel cognitively inspired hybrid framework, XGB–ANN–Attn, for interpretable and scalable retail analytics.MethodsThe model introduces feature-level fusion by integrating XGBoost-derived leaf embeddings with deep neural representations, enabling joint modeling of low-order statistical dependencies and high-order nonlinear feature interactions. A lightweight attention mechanism dynamically assigns instance-specific feature importance, enhancing both predictive performance and interpretability while maintaining linear computational complexity with respect to feature dimensionality. Unlike conventional ensemble approaches that operate at the decision level, the proposed framework integrates representations at the level of representations, reducing redundancy and improving computational efficiency. The model is designed for scalability, combining the log-linear complexity of gradient boosting with the linear scaling properties of neural networks and attention mechanisms.ResultsExperimental evaluation on the BigMart and Walmart datasets demonstrates superior performance compared to state-of-the-art models, achieving RMSE = 0.1584 and R2 = 0.9946 on BigMart, and RMSE = 0.8652 and R2 = 0.9568 on Walmart. Furthermore, the framework supports parallelization and distributed deployment, making it suitable for large-scale retail systems.DiscussionThe alignment between attention weights and SHAP explanations provides transparent and actionable insights. The results confirm that the proposed approach offers a scalable, interpretable, and high-performance solution for AI-driven decision support systems in big-data retail environments.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1812391</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1812391</link>
        <title><![CDATA[Interviewer effects on underreporting of traditional contraceptive use: evidence from National Family Health Survey-5]]></title>
        <pubdate>2026-08-20T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Rupalee Singh Chauhan</author><author>Laxmi Kant Dwivedi</author><author>Priyanka Dixit</author><author>Shiva S. Halli</author>
        <description><![CDATA[BackgroundSocioeconomic alignment between interviewers and respondents may play a critical role in shaping the accuracy of reported reproductive histories. This study investigates how interviewer characteristics contribute to the reporting of adoption of traditional contraceptive methods using reproductive event calendar data from the fifth round of the National Family Health Survey (2019–21) in India.Data and methodsUsing contraceptive calendar data, a cross-classified multilevel model was employed to assess the interviewer's influence on the reporting of traditional contraceptive method use. The analysis included women aged 15–49 years who reported contraceptive use in the calendar month(s), and the interviewers interviewed them.ResultsMultilevel cross-classified models show that interviewer characteristics significantly influence reporting of traditional contraceptive use following periods of non-use, childbirth, or pregnancy termination. Women interviewed by older interviewers (≥30 years; adjusted odds ratio (AOR) = 1.40, p < 0.01) and those who shared the same religion (AOR = 1.21, p ≤ 0.01) and marital status (AOR = 1.42, p < 0.01) as the interviewer were more likely to report traditional contraceptive method use. While women of the same age group (AOR = 0.89, p < 0.05), same-state residence (AOR = 0.82, p < 0.01), and same caste (AOR = 0.94, p < 0.01) as the interviewer were less likely to report traditional contraceptive method use. Older interviewers (AOR = 0.86, p < 0.01), those with secondary education (AOR = 0.54, p < 0.01), those belonging to the Muslim religion (AOR = 0.43, p < 0.01), and those with a lack of prior experience in conducting surveys (AOR = 0.90, p < 0.10) were associated with lower reporting. Moderate (AOR = 1.35, p < 0.01) and high (AOR = 1.22, p < 0.05) levels of prior experience with the same survey schedule in conducting the survey significantly improved reporting. Distribution of variance across different levels indicates substantial variation at the interviewer level (28.50%), followed by the district (18.50%) and PSU (13.46%) levels, suggesting that interviewer effects play a significant role in shaping the outcome.Discussion and conclusionInterviewer characteristics such as education, age, and prior survey experience influence the reporting of traditional contraceptive use. Similarities or differences between interviewers and respondents in characteristics such as age, marital status, religion, caste, and state of residence also affect the accuracy and completeness of reporting traditional contraceptive use in contraceptive calendar data.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1927115</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1927115</link>
        <title><![CDATA[A novel two-parameter discrete exponentiated exponential distribution with its neutrosophic extension: mathematical characterization and real-world count data modeling]]></title>
        <pubdate>2026-08-18T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Furqan Ahmed P.</author><author>Sujatha V.</author>
        <description><![CDATA[This paper proposes and studies the Discrete Exponentiated Exponential (DEE) distribution, a new two-parameter discrete lifetime model obtained by applying the survival discretization technique to the exponentiated exponential distribution. The DEE distribution possesses a key structural advantage: the single shape parameter α simultaneously controls both the hazard rate shape (increasing for 0 < α < 1, decreasing for α > 1) and the dispersion regime (over-dispersion for small α, under-dispersion for large α), a dual flexibility that most competing discrete models cannot replicate without structural modification. These properties are rigorously established in closed form via the Glaser technique and numerical investigation of the dispersion index. Closed-form expressions are derived for the probability mass function, cumulative distribution function, survival and hazard rate functions, moment generating and characteristic functions, raw and central moments up to the fourth order, and the mean residual life function. Parameter estimation is carried out via maximum likelihood, and a comprehensive Monte Carlo simulation study assesses the finite-sample performance of the estimators and benchmarks it against four competing discrete models under identical settings. The practical advantage of the DEE distribution is demonstrated through three real-world count datasets from oncology, education, and nephrology; the DEE model exhibits competitive and frequently superior performance relative to seven benchmark discrete distributions. Finally, a neutrosophic generalization (DNEE) extends the framework to count data characterized by vagueness, incompleteness, or indeterminacy, with accompanying estimation and illustrative numerical analysis.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1884673</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1884673</link>
        <title><![CDATA[Benchmarking retrieval augmented generation LLMs for Arabic noise robustness]]></title>
        <pubdate>2026-08-17T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Saleh Almohaimeed</author><author>Abdulrahman Alabduljabbar</author><author>Mousa Jari</author><author>Mohammed Alkhowaiter</author><author>Saad Almohaimeed</author><author>Mohamad Mahmoud Al Rahhal</author>
        <description><![CDATA[Hallucination has become a serious concern in large language models (LLMs), as these models can generate useful yet incorrect or misleading information, which has led to growing research interest in retrieval-augmented generation (RAG) as a mitigation approach. RAG provides LLMs with access to external information, such as databases or documents, which can help them to answer users' questions. Currently, several benchmarks released measure RAG performance on various LLMs; however, evaluations of the noise robustness ability in Arabic are absent. In this paper, we systematically investigated the capabilities of state-of-the-art multilingual LLMs with regard to two essential RAG abilities, noise robustness and negative rejection. To accomplish this, we generated an Arabic benchmark consisting of 300 questions along with 6,196 documents. Then, we assessed the performance of six LLMs in relation to the two aforementioned RAG abilities. The results reveal that all six LLMs were negatively affected when the noise ratio in the external documents was increased. Under the highest noise settings at 80%, the best LLM performance was for Claude-4 sonnet, in which their performance decreased by only 4.67 percentage points. Furthermore, when it comes to the negative rejection task, there has been a significant impact on all six models. The best two models, Claude-4 sonnet and Llama-4, scored 90.67% and 85.67%, respectively, while smaller models, like GPT-3.5, scored 69.33%. Furthermore, our manual analysis reveals that many errors made by LLMs are attributed to over-caution behavior. LLMs often decline to respond probably due to training mechanisms designed to reduce hallucinations. Additionally, other errors occur when there is a high lexical similarity between the question and the words of noisy documents, which causes the model to rely on irrelevant content instead of the correct information.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1811835</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1811835</link>
        <title><![CDATA[Real-time AI-driven trend analytics for smart city digital services using streaming data]]></title>
        <pubdate>2026-08-14T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Zhuldyz Kalpeyeva</author><author>Abdul Razaque</author><author>Raissa Uskenbayeva</author><author>Aliya Beishenaly</author><author>Venera Elle</author><author>Aizhan Kassymova</author>
        <description><![CDATA[Real-time data analysis plays an important role in the operation of digital urban systems, where the behavior of residents and the load on services can change over short periods of time. At the same time, traditional analytical approaches based on batch data processing often do not allow timely detection of such changes, which leads to delayed and not always accurate management decisions. This paper introduces an artificial intelligence-driven smart city system (AISSC) for real-time data trend analysis. The proposed AISSC framework processes streaming data in a smart city digital environment. The proposed approach encompasses real-time feature generation and statistical techniques for identifying significant changes. The proposed AISSC solution is executed on the Python platform. The results demonstrate that the proposed AISSC solution achieves a directional accuracy of 98.6% for trend prediction, together with strong numerical forecasting performance with a MAPE of 6.3% and a WMAPE of 7.8%. The framework detects statistically significant deviations within 2.4 s at the sliding-window level while maintaining an end-to-end system update cycle of approximately 5 min. These results demonstrate the capability of the proposed framework to support reliable real-time trend analysis and short-term forecasting for smart city decision-making. This shows that the framework can reliably identify trends and accurately estimate demand for real-time smart city decision-making. The results demonstrate consistent performance improvements compared to representative baseline methods under identical streaming and computational constraints. The proposed AISSC facilitates real-time observation of urban dynamics and enhances short-term forecasting accuracy for informed decision-making in smart city management systems.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1811967</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1811967</link>
        <title><![CDATA[Enhancing IoT botnet detection with explainable ensemble learning]]></title>
        <pubdate>2026-08-12T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Linda Joseph</author><author>Sambath M.</author><author>Vivekanandan M.</author>
        <description><![CDATA[IntroductionInternet of Things (IoT) botnet detection faces significant challenges due to the growing intricacy and decreased transparency of Machine Learning (ML) models.MethodsIn this work, we provide an ensemble-based detection system that makes use of a voting classifier made up of a Boosted Decision Tree and a Bagged Random Forest. Using manual feature extraction, the model is trained and assessed using the N-BaIoT dataset. The popular Explainable AI (XAI) method, SHapley Additive exPlanations (SHAP), is used to analyze feature contributions across models while taking important factors like consistency, sensitivity, and monotonicity into account in order to improve the interpretability of the features. SHAP eases the black-box characteristic of sophisticated machine learning models by quantifying the influence of specific features, hence facilitating transparent model interpretation.ResultsAccording to experimental data, the ensemble model is more sensitive than individual classifiers. Additionally, dynamic changes in SHAP values are shown by the weight adjustments made within the voting classifier, highlighting the impact of weight tuning on feature importance.DiscussionThis study highlights the effectiveness of integrating SHAP-based XAI into ensembled models, enhancing the transparency, interpretability, and reliability of IoT botnet detection systems.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1887075</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1887075</link>
        <title><![CDATA[DKFraudNet: a knowledge-guided adversarial learning framework for fraud user detection]]></title>
        <pubdate>2026-08-12T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Yingjun Shen</author><author>Renda Shi</author><author>Kaixi Song</author><author>Yunpeng Li</author>
        <description><![CDATA[IntroductionFraud user identification in telecommunications is hindered by scarce, noisy, and imbalanced labels, while expert rules may provide ambiguous or contradictory evidence.MethodsWe propose DKFraudNet, a knowledge-guided framework that integrates domain knowledge regularization, an attention-adaptive conditional generative adversarial network, and virtual category learning. Expert rules are organized into a Deterministic-Ambiguous-Contradictory evidence taxonomy for controlled pseudo-labeling. Class-conditioned augmentation alleviates data imbalance, and kernel-based similarity refines ambiguous samples. The framework was evaluated using two real-world city-level telecommunications datasets containing 59,015 and 52,183 users, respectively.ResultsDKFraudNet consistently improved downstream classification performance under weak supervision. On the independent City 2 validation set, the CatBoost instantiation achieved an accuracy of 0.911 and an F1-score of 0.909. For operational review prioritization, DKFraudNet-XGBoost required reviewing 40.2% of unverified users to capture 80% of fraud users, corresponding to a 49.71% workload reduction relative to random review, with an AUPRC of 0.954. In operator-side blind verification, genuine-member identification and false-member screening achieved accuracies of 98.9% and 93.3%, respectively.DiscussionThe results show that combining structured domain evidence, adaptive generative augmentation, and uncertainty-aware refinement improves robustness, data efficiency, and operational usefulness for fraud detection under weak supervision.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1916523</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1916523</link>
        <title><![CDATA[Differentiated memory and scientific cognition in AI research agents]]></title>
        <pubdate>2026-08-05T00:00:00Z</pubdate>
        <category>Hypothesis and Theory</category>
        <author>Diego F. Cuadros</author><author>Abdoul-Aziz Maiga</author><author>Sid Thatham</author><author>Alvaro Ortiz</author><author>Margaret Powers-Fletcher</author><author>Ming Tang</author>
        <description><![CDATA[AI research agents increasingly support ideation, literature search, coding, experimental execution, analysis, and manuscript drafting across the scientific workflow. This progress advances automated discovery, but workflow automation is not scientific cognition. Scientific reasoning is cumulative and path-dependent: it depends on what a researcher has written, read, learned from critique, absorbed through experience, and used as habitual standards for judging novelty, rigor, feasibility, and significance. We propose Mnemo as a framework for modeling scientific cognition in AI research agents. First, scientific cognition may require differentiated memory, organized into distinct spaces for authored work, external reference, critique, experience, and judgment. Second, provenance should be treated not as passive metadata but as memory routing, because source origin helps determine cognitive function in reasoning. Third, new ideas may be better modeled as controlled collisions across memory spaces, filtered by judgment, rejection, and epistemic calibration, than as generic recombination from model priors. Mnemo motivates a research agenda for AI in science centered on routing fidelity, critique use, judgment alignment, rejection quality, and epistemic calibration.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1899143</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1899143</link>
        <title><![CDATA[A computational analysis of human rights discourse in news media using institutional theory]]></title>
        <pubdate>2026-07-31T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Joel Andrew B. Cruz Jr</author><author>Ma. Rowena R. Caguiat</author>
        <description><![CDATA[IntroductionNews coverage of politically sensitive events involves framing choices that accumulate across thousands of headlines, years, and media institutions. Likewise, no single reader or journalist can track these patterns as they happen. This study examines how English-language news media framed Rodrigo Duterte's anti-drug campaign in the Philippines from June 2016 to June 2025.MethodsThe study applies BERTopic topic modeling and VADER sentiment analysis to 1,133 headlines collected through Google News's RSS endpoint. Institutional Theory organizes the results around three pressures: normative, coercive, and mimetic.ResultsTen topic clusters emerged from the corpus, centered on human rights violations, International Criminal Court (ICC) legal proceedings, and post-presidency accountability narratives. Domestic Philippine outlets covered human rights themes in a larger share of headlines than international outlets during the campaign's early years, but that pattern reversed after Duterte left office in 2022. By 2024 and 2025, international outlets covered human rights themes in roughly 64 percent of headlines, compared with about half for domestic outlets. Sentiment analysis classified 807 headlines as negative, 203 as neutral, and 123 as positive. Negative coverage rose sharply in early 2025, around Duterte's arrest on an ICC warrant. Headline language also clustered around three institutional events: the 2018 ICC filing, the 2020 United Nations (UN) Human Rights Report, and the 2021 ICC investigation authorization, each marked by distinct word patterns across positive, negative, and neutral coverage.DiscussionNormative pressure explains why domestic outlets led on human rights framing before any international body had acted. Coercive pressure from the ICC and UN corresponds to the sentiment spikes around each event. Mimetic pressure explains how framing conventions moved between domestic and international outlets over time, first from domestic to international sources and later in the other direction. The study also considers how Google News's own curation choices shape which outlets and frames appear in the corpus. These findings suggest computational analysis is both a research method and a way to hold a media record open to scrutiny after the news cycle moves on.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1886902</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1886902</link>
        <title><![CDATA[Shifting supply chain interdependencies among global automakers]]></title>
        <pubdate>2026-07-23T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Hiromitsu Goto</author><author>Wataru Souma</author>
        <description><![CDATA[Against the backdrop of electrification and supply chain resilience, the global automotive industry is in an era of great transformation. Using component supply information provided by MarkLines, this study investigated the impact on inter-firm dependencies among major global automakers from 2018 to 2024. Specifically, it analyzed changes in the community structure of the interdependency network between manufacturers, based on common suppliers for each model year and component category. The results revealed the impact of electrification: a reduction of roughly two-thirds (about 64%) in transactions for internal combustion engine (ICE) powertrains and an approximately eight-fold increase in e-powertrain transactions. Furthermore, the results of community detection also revealed a structural reorganization: while geographical clustering among manufacturers intensified for ICE components, new cross-border interdependencies formed for e-powertrain components, leading to the fragmentation of the traditionally integrated Japan-US-Europe bloc and the emergence of a China-centric ecosystem. This study provides new empirical evidence on the structural realignment of the global automotive value chain, offering important implications for management strategy and industrial policy in an era of great technological and geopolitical change.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1846964</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1846964</link>
        <title><![CDATA[AI applied to Saudi Arabia higher education: systematic literature review—SLR with PRISMA and VOSviewer]]></title>
        <pubdate>2026-07-22T00:00:00Z</pubdate>
        <category>Systematic Review</category>
        <author>Jehad Alqurni</author>
        <description><![CDATA[This study explored the role of artificial intelligence in higher education in Saudi Arabia. It aims to identify trends, key contributors, influential papers, collaborations, and important research areas from January 2022 to February 2026 to guide future research. The researchers used bibliometric and content analyses. This combined quantitative descriptive methods and network analysis with qualitative content analysis of the most-cited articles. They extracted data from Scopus and the Web of Science, resulting in 66 documents after removing duplicates, editorials, and notes. The analytical techniques included the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA), co-word analysis, citation analysis, co-authorship analysis, and bibliographic coupling. VOSviewer supported the visualization. The key findings show that King Abdulaziz University, Qassim University, and King Saud University are the top contributors. A total of 52 papers have been published in journals indexed by WOS and Scopus, while 14 papers are indexed only by Scopus. Additionally, five major groups revealed significant correlations among various word pairs: High-education, Artificial intelligence, AI Chatbot, learning systems, and ChatGPT. The leading journals include Sustainability, Acta Psychological, and Systems, with notable authors such as Al-Harbi. Content analysis highlights the potential of AI to improve learning, boost administrative efficiency, and drive innovation in Saudi Arabia. This study offers practical recommendations for students, universities, and policy makers. This study has several limitations that should be addressed in future studies.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1878242</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1878242</link>
        <title><![CDATA[A cost-sensitive random forest framework for ARP spoofing detection in Internet of Medical Things networks]]></title>
        <pubdate>2026-07-22T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Siddhartha Singhal</author><author>Kakelli Anil Kumar</author>
        <description><![CDATA[IntroductionARP spoofing poses a major security threat to Internet of Medical Things (IoMT) networks by enabling man-in-the-middle attacks that compromise the integrity of life-critical communications. Existing intrusion detection methods fail to simultaneously address temporal attack dynamics, unequal medical safety requirements, and explicit control of false negative rates.MethodsThis study proposes the Self-Healing IoT-Optimized Random Forest (SH-IORF) framework, which integrates temporal behavioral feature engineering, validation-guided cost-sensitive learning, and medical safety-constrained threshold optimization. To ensure methodological rigor and prevent information leakage, a stratified three-way partitioning strategy consisting of training, validation, and completely held-out testing datasets was employed. Class penalty weights and operating thresholds were determined exclusively from the validation dataset.ResultsExperimental evaluation on the CICIoMT2024 benchmark demonstrated that the proposed SH-IORF framework achieved 99.90% accuracy, 99.83% recall, 99.95% precision, a 0.9989 F1-score, and an AUC-ROC of 0.9996. The framework limited the false negative rate to 0.17%, satisfying the predefined medical safety constraint (FNR ≤ 0.5%), corresponding to 40 missed detections among 23,390 attack samples and 12 false alarms across 28,768 benign traffic instances.DiscussionThe results demonstrate that the proposed framework provides stable and safety-oriented intrusion detection capability under heterogeneous IoMT deployment conditions while maintaining strict testing independence and robust performance under rigorous evaluation settings.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1832790</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1832790</link>
        <title><![CDATA[Machine learning approach for predicting the severity risk of obstructive sleep apnea syndrome]]></title>
        <pubdate>2026-07-16T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Qi Wang</author><author>Xiaoyu Yang</author><author>Shuran Xu</author><author>Haohao Wu</author><author>Guixuan Wang</author><author>Huixian Liu</author><author>Ronghua Chen</author><author>Fengming Xu</author><author>Cheng Wang</author><author>Kang Du</author>
        <description><![CDATA[BackgroundObstructive Sleep Apnea-Hypopnea Syndrome (OSAHS) has a high global prevalence and is prone to causing various serious complications. Our objective is to develop severity stratification of OSAHS by integrating multiple commonly available clinical features based on machine learning (ML).Materials and methodsThis study collected data from 432 cases at Qujing Central Hospital in Yunnan Province, integrating 25 clinical feature variables. The cases were randomly split into training (70%) and validation (30%) sets. The importance of the 25 features was analyzed.ResultsIt showed that the HCY, TBIL, BMI, GGT, and Age made significant contributions to OSAHS severity. We established five machine learning models-Multilayer Perceptron (MLP), Random Forest, XGBoost, LightGBM, and Support Vector Machine (SVM)-by integrating 25 clinical features. Through cross-validation and continuous adjustment of model parameters, the optimal predictive model was determined. By calculating model accuracy and F1-score, XGBoost was identified as the best-performing model, achieving an area under the curve (AUC) of 0.63, an accuracy of 75% and an F1-score of 65.60.ConclusionIn this study, we established a predictive model for the severity stratification of OSAHS based on machine learning algorithms. The XGBoost model demonstrated superior predictive performance.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1878260</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1878260</link>
        <title><![CDATA[A dataset-centric review of IoT and IIoT intrusion detection: realism, evaluation biases, and future research directions]]></title>
        <pubdate>2026-07-10T00:00:00Z</pubdate>
        <category>Review</category>
        <author>Dwarsala Sreedhar Reddy</author><author>Kakelli Anil Kumar</author>
        <description><![CDATA[The rapid growth of IoT and IIoT expands the cyber-attack surface of interconnected and safety-critical systems, and, as such, IDSs have become a fundamental security mechanism. Although very impressive results have been reported for machine learning and deep learning-based IDS in benchmark datasets, these gains often do not generalize to real-world deployments owing to dataset design limitations, realism deficits, and evaluation biases, rather than inherent flaws in detection algorithms, which can lead to significant vulnerabilities in actual operational environments. This study presents a dataset-centric review of widely used intrusion detection datasets from the IIoT, IoT, and traditional network domains. A unified taxonomy differentiates datasets based on the domain context, traffic representation, protocol semantics, and attack modeling assumptions. Based on a common analytical framework, each dataset was reviewed regarding its realism, coverage of the threats, class imbalance, temporal continuity, and modern ML/DL-based evaluation of the IDS. The cross-dataset analysis conducted in this study shows that, in addition to the fact that model architecture and feature engineering play a major role, several studies indicate that the simplicity of the datasets, the class imbalance, and the repetitive attack patterns as well as the evaluation methods can affect accuracy of the IDS. This work further underlines the remaining gaps, such as zero-day and adaptive attacks, limited encrypted traffic, weak temporal evolution, poor support for federated learning, and sparse annotations for explainable IDSs. Finally, this study presents future directions for dataset design aligned with the requirements of next-generation IDSs by highlighting digital twin-based IIoT environments, edge-cloud collaborative data generation, sequential traffic modeling, and explainability-oriented annotations that can ensure robust, trustworthy, and deployment-ready IDS solutions.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1728498</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1728498</link>
        <title><![CDATA[A prognostic tool for pulmonary collapse: nomogram-based prediction of 28-day mortality]]></title>
        <pubdate>2026-07-09T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Xinming He</author><author>Wenchong Yu</author><author>Yuling Li</author><author>Ao Ma</author><author>Zhichao Meng</author><author>Jiehao Zhu</author><author>Minghui Tan</author><author>Xiaodong Zhao</author><author>Mu Chen</author>
        <description><![CDATA[BackgroundPulmonary collapse is a common and serious respiratory condition, but there is no dedicated bedside tool to estimate prognosis. This study aimed to develop a nomogram to predict 28-day mortality in patients with pulmonary collapse.MethodsWe extracted data for patients with pulmonary collapse from MIMIC-III, identified predictors using regression analyses, and used MIMIC-IV for temporal validation. We then built a nomogram based on the selected predictors. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC), AUC comparisons using the DeLong test, reclassification (NRI and IDI), calibration (calibration curves, calibration slope, and Brier score), and decision curve analysis (DCA).ResultsA total of 4,088 patients with pulmonary collapse were included in the study. Logistic regression analysis identified twelve independent predictive factors associated with 28-day mortality: age (OR = 1.01, P =0.040), married status (OR =0 .63, P =0 .040), Glasgow Coma Scale score (OR = 0.94, P = 0.02), creatinine (OR = 0.82, P = 0.03), chloride ions (OR = 0.85, P = 0.03), sodium ions (OR = 1.19, P = 0.02), blood urea nitrogen (OR = 1.02, P < 0.001), white blood cell count (OR = 1.04, P < 0.001), heart rate (OR = 1.02, P = 0.02), respiratory rate (OR = 1.05, P = 0.03), temperature (OR = 0.54, P < 0.001), and metastatic cancer (OR = 6.66, P < 0.001). The nomogram showed moderate discrimination and consistently higher AUC than Age+Gender, SOFA, and SAPSII across cohorts.ConclusionThis study identified factors associated with 28-day mortality in patients with pulmonary collapse and developed a nomogram for early risk stratification.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1837706</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1837706</link>
        <title><![CDATA[Noise-robust temporal–spectral fusion transformers for EEG-based cognitive state classification in aviation environments]]></title>
        <pubdate>2026-07-09T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Quynh Anh Nguyen</author><author>Nam Anh Dao</author><author>Long Nguyen</author>
        <description><![CDATA[Attention-related Pilot Performance Decrements (APPD) contribute substantially to aviation incidents, yet existing electroencephalography (EEG)-based monitoring methods often lack generalization, robustness to noise, and effective temporal–spectral integration. We propose a temporal–spectral fusion transformer (TF-T) combining multi-scale preprocessing, dual-stream temporal and spectral feature extraction, and transformer-based fusion with enhanced temporal–spectral integration and multi-resolution feature processing for multiclass cognitive state recognition. Three variants (TF-T1–TF-T3) are evaluated on controlled and ecologically realistic EEG datasets under clean and noise-augmented (Gaussian, Uniform, COMBO) conditions, using chronological partitioning to avoid temporal leakage. TF-T2 achieves the highest clean-data accuracy (99.2%), while TF-T3 offers superior robustness, improving Macro-F1 by ~4.5–4.7 points across all noise types and outperforming state-of-the-art baselines by up to +8 Macro-F1 under COMBO noise, supporting its deployment in perturbation-prone aviation environments.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1842233</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1842233</link>
        <title><![CDATA[Evolutionary multi-agent reinforcement learning for crisis-aware demographic policy optimization]]></title>
        <pubdate>2026-07-08T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Anton V. Dozhdikov</author><author>Arseniy M. Sitkovskiy</author>
        <description><![CDATA[Demographic systems face unprecedented challenges from simultaneous crises. Conventional statistical demography techniques and agent–based models often struggle to capture nonlinear inter–regional interactions during periods of severe socio–economic disruption. To address this, we propose MADDPG–EVO–DGM, a hybrid algorithm that integrates multi–agent deep reinforcement learning with evolutionary optimisation and meta–learning principles to model regional demographic processes under multiple crisis scenarios. Each region is treated as an autonomous agent learning to steer demographic policy levers, while periodic evolutionary “boosters” overcome local optima via population–based perturbations of actor network parameters. Additionally, a Darwin–Gödel Machine–inspired meta–learning mechanism adapts the booster triggers, enabling self–improvement in the learning process. We evaluate MADDPG–EVO–DGM on a simulation environment calibrated with real demographic data for eight federal regions of the Russian Federation over the period 2000–2024 and subject to ten concurrent crisis scenarios (e.g., pandemic, geopolitical conflict, economic collapse). Experiments demonstrate significantly faster convergence and improved performance over a baseline MADDPG: the hybrid approach achieves a higher final average reward (252.57 vs. 243.07) and 3.4 × lower convergence variance (σ = 0.24 vs. 0.80), indicating more reliable training. It also exhibits qualitative performance jumps of +68% during evolutionary phases and maintains 35%–45% greater resilience under crisis shocks compared to the baseline. To our knowledge, this is the first application of multi–agent reinforcement learning to large–scale demographic modeling under crises, opening new possibilities for evidence–based, crisis–resilient population policy design. Code, data, and logs are provided to ensure reproducibility.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1883452</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1883452</link>
        <title><![CDATA[Cross-model evaluation of phishing detectors against LLM-generated emails]]></title>
        <pubdate>2026-07-07T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Rommel Gutierrez</author><author>William Villegas-Ch</author><author>Jaime Govea</author>
        <description><![CDATA[Phishing remains a prevalent cyberattack vector, and the widespread adoption of large language models (LLMs) has enabled adversaries to generate grammatically correct and contextually coherent phishing emails at scale, against which conventional detection systems are less effective. Although stylometric methods achieve over 95% accuracy within a single generator, their performance has not been systematically evaluated when the source model changes between training and deployment. This represents a significant gap, as adversaries can switch generators rapidly. A balanced corpus of 9,986 phishing emails was assembled, comprising 4,986 emails generated by three modern LLMs (GPT-4.1, DeepSeek 3.2, and Llama 3.3 70B) across five thematic categories, and 5,000 human phishing emails sampled in a stratified manner from five public sources. Seventeen stylometric features were extracted, and Logistic Regression and XGBoost classifiers were evaluated under intra-model, cross-model, threshold-recalibrated, cross-dataset, and aggregated-pool settings. Intra-model F1 scores reached 0.96 under stratified cross-validation and 0.999 on held-out splits used for the cross-model matrix. However, cross-model F1 dropped by 28.0 percentage points under the default decision threshold of 0.5. Notably, the area under the receiver operating characteristic curve remained above 0.96 in every off-diagonal cell, indicating that discriminative information is preserved even though the decision threshold is generator-specific. Recalibrating the threshold on a small target subset reduced the gap to 4.0 percentage points (an 86% reduction), and an aggregated-pool detector achieved F1 = 0.997 on each generator. This work reframes cross-model phishing detection from a problem of model incompatibility to one of practical calibration, and provides two deployable solutions, threshold recalibration on a small target slice and aggregated-pool training, along with a publicly released multi-LLM corpus.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1807559</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1807559</link>
        <title><![CDATA[Influence of localized roadway surface obstacles on vehicular emissions under real-world urban driving conditions]]></title>
        <pubdate>2026-06-23T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Victor Cardoso Oliveira</author><author>Thiago Iachiley Araújo de Souza</author><author>Nicole Souza Batista</author><author>Bruno Vieira Bertoncini</author><author>Verônica Teixeira Franco Castelo Branco</author>
        <description><![CDATA[IntroductionVehicular emissions are a major source of air pollution in tropical urban environments. While the impacts of technology, traffic flow, and driving behavior on pollutant formation are well established, the influence of pavement surface remains insufficiently understood. Pavement defects such as potholes, cracks, and depressions disturb vehicle operation and may increase real-world emissions. This study evaluates the influence of pavement obstacles on emissions of CO2, CO, and NOx across five urban road segments in Fortaleza, Brazil.MethodsA portable emissions measurement system (PEMS) collected second-by-second exhaust data during real driving. Roadway surface obstacles were captured through windshield-mounted images acquired at 1 Hz. A broader image pool comprising 44,175 roadway images was assembled from multiple urban roads for model development. From this pool, a final annotated dataset comprising 1,812 images with 3,158 labeled obstacles was used for training, validation, and testing the YOLOv8n detector, which was then applied to the five monitored road sections used in the synchronized emission analysis. Emission data and obstacle locations were synchronized, enabling comparison of pollutant rates along obstacle-present and obstacle-free segments, with emphasis on features likely to influence short-term driving behavior. Detection performance was evaluated using precision, recall, and mAP metrics.ResultsRoad segments with higher obstacle occurrence presented elevated emission rates. In the full dataset, maximum values reached 478.2 g/km for CO2, 491.96 mg/km for CO, and 100.266 mg/km for NOx. A filtered analysis excluding curves, intersection buffers, and visible traffic or pedestrian interference showed that obstacle-present observations still exhibited higher emissions than obstacle-free observations, with average increases of 26% for CO2, 31% for NOx, and 42% for CO. Spatial mapping showed that emission hotspots tended to occur in areas with frequent roadway surface obstacles and operational disturbances.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1871346</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1871346</link>
        <title><![CDATA[Adaptive class-aware feature selection for high-dimensional and imbalanced multi-class network intrusion detection]]></title>
        <pubdate>2026-06-23T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Joseph P. Mchina</author><author>Neema Mduma</author><author>Ramadhani S. Sinde</author>
        <description><![CDATA[High-dimensional feature spaces and severe class imbalance remain fundamental challenges for Machine Learning-based Network Intrusion Detection Systems (ML-NIDS), where minority attack categories are frequently overlooked during feature selection. Existing feature selection approaches commonly rely on global feature relevance measures and manually specified feature counts, which favor majority traffic classes and reduce sensitivity toward rare but critical attack categories. To address these limitations, this study proposes the Adaptive Class-Aware Feature Selection (ACAFS) framework for multi-class intrusion detection. Unlike conventional approaches, ACAFS introduces a data-driven adaptive feature count mechanism based on permutation null hypothesis testing, a Class-Aware Composite Mutual Information scoring strategy that explicitly preserves minority-class discriminative information, and a coordinated two-stage feature selection framework that combines statistical filtering with XGBoost-based model refinement. The framework was evaluated independently on the CSE-CIC-IDS2018 benchmark dataset and a Simulated University Network Environment (SUNE) dataset representing Tanzanian higher learning institution networks. Experimental results demonstrate that ACAFS substantially reduces feature dimensionality while improving balanced intrusion detection performance. On CSE-CIC-IDS2018, ACAFS reduced the feature space from 74 to 22 features, representing a 70.3% dimensionality reduction, while the Two-Stage CNN achieved 99.39% accuracy, 99.40% F1-score, and a false positive rate of 0.09%. The framework further achieved 98.59% recall for Web_Attacks despite severe class imbalance, demonstrating effective preservation of minority-class discriminative features. On the SUNE dataset, ACAFS independently selected 18 features and maintained stable detection performance without dataset-specific manual tuning, confirming its adaptability across heterogeneous network environments. These results confirm that adaptive and class-aware feature selection can simultaneously reduce feature redundancy, improve minority attack detection, and maintain robust intrusion detection performance across diverse network traffic environments.]]></description>
      </item>
      </channel>
    </rss>