<?xml version="1.0" encoding="utf-8"?>
    <rss version="2.0">
      <channel xmlns:content="http://purl.org/rss/1.0/modules/content/">
        <title>Frontiers in Big Data | Machine Learning and Artificial Intelligence section | New and Recent Articles</title>
        <link>https://www.frontiersin.org/journals/big-data/sections/machine-learning-and-artificial-intelligence</link>
        <description>RSS Feed for Machine Learning and Artificial Intelligence section in the Frontiers in Big Data journal | New and Recent Articles</description>
        <language>en-us</language>
        <generator>Frontiers Feed Generator,version:1</generator>
        <pubDate>2026-08-10T21:20:57.93+00:00</pubDate>
        <ttl>60</ttl>
        <item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1916523</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1916523</link>
        <title><![CDATA[Differentiated memory and scientific cognition in AI research agents]]></title>
        <pubdate>2026-08-05T00:00:00Z</pubdate>
        <category>Hypothesis and Theory</category>
        <author>Diego F. Cuadros</author><author>Abdoul-Aziz Maiga</author><author>Sid Thatham</author><author>Alvaro Ortiz</author><author>Margaret Powers-Fletcher</author><author>Ming Tang</author>
        <description><![CDATA[AI research agents increasingly support ideation, literature search, coding, experimental execution, analysis, and manuscript drafting across the scientific workflow. This progress advances automated discovery, but workflow automation is not scientific cognition. Scientific reasoning is cumulative and path-dependent: it depends on what a researcher has written, read, learned from critique, absorbed through experience, and used as habitual standards for judging novelty, rigor, feasibility, and significance. We propose Mnemo as a framework for modeling scientific cognition in AI research agents. First, scientific cognition may require differentiated memory, organized into distinct spaces for authored work, external reference, critique, experience, and judgment. Second, provenance should be treated not as passive metadata but as memory routing, because source origin helps determine cognitive function in reasoning. Third, new ideas may be better modeled as controlled collisions across memory spaces, filtered by judgment, rejection, and epistemic calibration, than as generic recombination from model priors. Mnemo motivates a research agenda for AI in science centered on routing fidelity, critique use, judgment alignment, rejection quality, and epistemic calibration.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1832790</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1832790</link>
        <title><![CDATA[Machine learning approach for predicting the severity risk of obstructive sleep apnea syndrome]]></title>
        <pubdate>2026-07-16T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Qi Wang</author><author>Xiaoyu Yang</author><author>Shuran Xu</author><author>Haohao Wu</author><author>Guixuan Wang</author><author>Huixian Liu</author><author>Ronghua Chen</author><author>Fengming Xu</author><author>Cheng Wang</author><author>Kang Du</author>
        <description><![CDATA[BackgroundObstructive Sleep Apnea-Hypopnea Syndrome (OSAHS) has a high global prevalence and is prone to causing various serious complications. Our objective is to develop severity stratification of OSAHS by integrating multiple commonly available clinical features based on machine learning (ML).Materials and methodsThis study collected data from 432 cases at Qujing Central Hospital in Yunnan Province, integrating 25 clinical feature variables. The cases were randomly split into training (70%) and validation (30%) sets. The importance of the 25 features was analyzed.ResultsIt showed that the HCY, TBIL, BMI, GGT, and Age made significant contributions to OSAHS severity. We established five machine learning models-Multilayer Perceptron (MLP), Random Forest, XGBoost, LightGBM, and Support Vector Machine (SVM)-by integrating 25 clinical features. Through cross-validation and continuous adjustment of model parameters, the optimal predictive model was determined. By calculating model accuracy and F1-score, XGBoost was identified as the best-performing model, achieving an area under the curve (AUC) of 0.63, an accuracy of 75% and an F1-score of 65.60.ConclusionIn this study, we established a predictive model for the severity stratification of OSAHS based on machine learning algorithms. The XGBoost model demonstrated superior predictive performance.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1837706</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1837706</link>
        <title><![CDATA[Noise-robust temporal–spectral fusion transformers for EEG-based cognitive state classification in aviation environments]]></title>
        <pubdate>2026-07-09T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Quynh Anh Nguyen</author><author>Nam Anh Dao</author><author>Long Nguyen</author>
        <description><![CDATA[Attention-related Pilot Performance Decrements (APPD) contribute substantially to aviation incidents, yet existing electroencephalography (EEG)-based monitoring methods often lack generalization, robustness to noise, and effective temporal–spectral integration. We propose a temporal–spectral fusion transformer (TF-T) combining multi-scale preprocessing, dual-stream temporal and spectral feature extraction, and transformer-based fusion with enhanced temporal–spectral integration and multi-resolution feature processing for multiclass cognitive state recognition. Three variants (TF-T1–TF-T3) are evaluated on controlled and ecologically realistic EEG datasets under clean and noise-augmented (Gaussian, Uniform, COMBO) conditions, using chronological partitioning to avoid temporal leakage. TF-T2 achieves the highest clean-data accuracy (99.2%), while TF-T3 offers superior robustness, improving Macro-F1 by ~4.5–4.7 points across all noise types and outperforming state-of-the-art baselines by up to +8 Macro-F1 under COMBO noise, supporting its deployment in perturbation-prone aviation environments.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1842233</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1842233</link>
        <title><![CDATA[Evolutionary multi-agent reinforcement learning for crisis-aware demographic policy optimization]]></title>
        <pubdate>2026-07-08T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Anton V. Dozhdikov</author><author>Arseniy M. Sitkovskiy</author>
        <description><![CDATA[Demographic systems face unprecedented challenges from simultaneous crises. Conventional statistical demography techniques and agent–based models often struggle to capture nonlinear inter–regional interactions during periods of severe socio–economic disruption. To address this, we propose MADDPG–EVO–DGM, a hybrid algorithm that integrates multi–agent deep reinforcement learning with evolutionary optimisation and meta–learning principles to model regional demographic processes under multiple crisis scenarios. Each region is treated as an autonomous agent learning to steer demographic policy levers, while periodic evolutionary “boosters” overcome local optima via population–based perturbations of actor network parameters. Additionally, a Darwin–Gödel Machine–inspired meta–learning mechanism adapts the booster triggers, enabling self–improvement in the learning process. We evaluate MADDPG–EVO–DGM on a simulation environment calibrated with real demographic data for eight federal regions of the Russian Federation over the period 2000–2024 and subject to ten concurrent crisis scenarios (e.g., pandemic, geopolitical conflict, economic collapse). Experiments demonstrate significantly faster convergence and improved performance over a baseline MADDPG: the hybrid approach achieves a higher final average reward (252.57 vs. 243.07) and 3.4 × lower convergence variance (σ = 0.24 vs. 0.80), indicating more reliable training. It also exhibits qualitative performance jumps of +68% during evolutionary phases and maintains 35%–45% greater resilience under crisis shocks compared to the baseline. To our knowledge, this is the first application of multi–agent reinforcement learning to large–scale demographic modeling under crises, opening new possibilities for evidence–based, crisis–resilient population policy design. Code, data, and logs are provided to ensure reproducibility.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1838191</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1838191</link>
        <title><![CDATA[Measuring the impact of virtualization and containerization on the environment when using GPUs for processing the AI models]]></title>
        <pubdate>2026-06-16T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Safaa Hriez</author><author>Mohammad Haikal</author>
        <description><![CDATA[IntroductionThe rapid growth of artificial intelligence (AI) has significantly increased computing demand, intensifying the operational strain on the computing environment. While virtualization and containerization are established technologies for resource optimization, their comparative energy efficiency and environmental impact, particularly under GPU-accelerated AI workloads, are not well-quantified.MethodsThis study evaluates the energy consumption and environmental impact of virtualization and containerization technologies when using Graphics Processing Units (GPUs) for AI model execution. Employing a computer vision benchmark, the performance, GPU resource utilization, and power consumption were measured. The experiment involved training a DenseNet-121 model on the MNIST dataset within a VirtualBox virtual machine and a Docker container environment.ResultsThe analysis indicates that containerization consistently surpasses virtualization in energy efficiency. Specifically, Docker container configuration demonstrated an approximately 21.6% reduction in total energy consumption, and a corresponding reduction in carbon dioxide (CO2) emissions compared to a VirtualBox virtual machine. Furthermore, containerization exhibited lower average and peak GPU utilization and power consumption.DiscussionThese findings demonstrate that containerization offers a more energy-efficient and environmentally sustainable approach than VirtualBox virtualization for the specific GPU-enabled AI workload evaluated in this study. Statistical significance testing indicates that the observed performance differentials are significant, supporting the validity of the results within the experimental scope of this work.ConclusionImplementing containerization in this experimental setup may reduce energy consumption and environmental impact without compromising computational performance. Future studies should extend these analyses to larger neural network models, diverse AI workloads, and heterogeneous GPU platforms to enhance the generalizability of these findings beyond the current single-system experimental configuration.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1752468</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1752468</link>
        <title><![CDATA[Data field theory: a geometric framework for learning on Riemannian manifolds with synthetic validation and limitation analysis]]></title>
        <pubdate>2026-06-11T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Mohammadreza Nehzati</author>
        <description><![CDATA[IntroductionConventional machine learning treats learning as parameter optimization, lacking a first-principles framework for phenomena like criticality, generalization, and causal structure. We introduce Data Field Theory (DFT), a mathematical framework modelling learning as the evolution of a data field governed by stochastic partial differential equations on Riemannian manifolds. This work aims to validate DFT's core predictions in settings where its geometric assumptions hold, while honestly assessing its empirical limitations.MethodsWe formulate learning as a field φ:M×ℝ≥0→ℝk evolving on a spherical manifold. To test DFT, we implement a hierarchical classification task using synthetic data drawn from von Mises-Fisher distributions, ensuring match with the manifold geometry. We derive four key predictions: (1) critical exponents near concept formation, (2) a spectral robustness law linking Eigen gaps to out-of-distribution (OOD) error, (3) finite-speed causal propagation from hyperbolic regularization, and (4) approximate rotational equivariance via a Ward identity. We also conduct a preliminary real-data experiment projecting MNIST digits onto the sphere.ResultsSynthetic experiments validate all four predictions: (1) Correlation length diverges as ξ(t)~|t-tc|-ν with ν = 0.63 ± 0.04, accompanied by 1/f fluctuations; (2) OOD generalization error scales as ϵOOD∝mgap-2 (ρ = −0.78, p < 10−6); (3) Causal propagation speed ceff = 0.98 ± 0.03 (theory maximum cmax = 1.0) under hyperbolic regularization; and (4) Ward identity residual R = 0.0032 ± 0.0008 converging as R∝h1.02. However, on real-world MNIST-sphere data, DFT achieves only 15.7% accuracy versus 51.7% for k-NN, revealing critical limitations.DiscussionDFT successfully predicts emergent phenomena criticality, spectral robustness, bounded causality, and approximate equivariance under ideal geometric conditions, supporting its theoretical validity. The poor real-data performance highlights key gaps: the current framework lacks adaptive metric learning, noise robustness, and hierarchical feature extraction present in real images. These results establish DFT as a principled mathematical foundation for learning as field dynamics while clearly delineating necessary extensions for practical applicability.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1883246</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1883246</link>
        <title><![CDATA[Correction: Explainable gradient convolutional vector fuzzy pattern analysis based on ensemble model for facial expression recognition]]></title>
        <pubdate>2026-06-10T00:00:00Z</pubdate>
        <category>Correction</category>
        <author>Lakshmi Sarvani Videla</author><author>Babu Reddy Mukamalla</author>
        <description></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1825213</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1825213</link>
        <title><![CDATA[When uncertainty guides learning: a highly effective approach to kidney disease classification in CT imaging]]></title>
        <pubdate>2026-06-09T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Muslima Akter</author><author>Fahmid Al Farid</author><author>Md Yousuf Ahmad</author><author>Md Azad Hossain Raju</author><author>Sowad Rahman</author><author>Jia Uddin</author><author>Hezerul Bin Abdul Karim</author>
        <description><![CDATA[The high cost of expert annotations significantly hinders the advancement of deep learning models for clinical medical imaging. This work introduces an efficient entropy-based active learning framework that achieves outstanding classification performance for renal abnormalities (Normal, Cyst, Stone, Tumor) in CT scans while requiring only a minimal amount of labeled data. The dataset comprises 12,446 CT slices split 70/15/15 into training (8,716), validation (1,865), and test (1,865) partitions via stratified sampling. Starting with only 200 randomly selected images and employing predictive entropy for uncertainty sampling on a pretrained ResNet-50 backbone, the proposed method attains 99.71% ± 0.25% mean test accuracy (95% CI: [99.30, 99.94]) across five independent runs after just six query cycles on the standard 12,446-image CT kidney dataset. Our method uses only 2,000 labeled training images, representing 22.9% of the 8,716-image training partition (a 77.1% reduction in required annotations relative to full supervision of the training set). This performance matches or exceeds prior fully supervised methods trained on the complete labeled training partition while demonstrating substantially improved sample efficiency, particularly in early annotation cycles where entropy-guided selection converges significantly faster than random sampling. Statistical testing across five repeated runs confirms that results are stable (Shapiro-Wilk p = 0.148). The framework exhibits exceptional sample efficiency as described by an empirically fitted power-law curve with a fitted exponent of 1.2, and empirically observed uncertainty decay with a rate of 0.92. These results offer both practical insights into annotation efficiency and substantial application value in the medical imaging domain.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1796969</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1796969</link>
        <title><![CDATA[TCMB: cross-model multi-level cross-attention network with Taylor-based loss for multimodal fake news detection]]></title>
        <pubdate>2026-06-05T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Santosh Kumar Banbhrani</author>
        <description><![CDATA[IntroductionThe rapid spread of misinformation across social media platforms, websites, and online communication channels has made fake news detection a critical task in the digital era. Although various computational approaches have been developed to identify fake news, many existing methods suffer from limitations such as biased training datasets and high rates of false positives and false negatives. To address these challenges, this study proposes a Multimodal Cross Attention Network with Taylor-based Cross Entropy Mean Bias (MMCN_TCMB) model for detecting multimodal fake news.MethodsThe proposed approach utilizes multimodal inputs consisting of textual and visual content obtained from fake news datasets. The textual information in news posts is first tokenized using Bidirectional Encoder Representations from Transformers (BERT). Feature extraction is then performed using Word2Vec and Term Frequency–Inverse Gravity Moment (TF-IGM). Simultaneously, images associated with news posts undergo preprocessing through Contrast Limited Adaptive Histogram Equalization and Histogram Equalization (CLAHE-HE), followed by feature extraction using ResNet. The extracted textual and visual features are combined and processed through the MMCN framework. The learning mechanism of the network is enhanced using the Taylor-based Cross Entropy Mean Bias (TCMB) loss function to improve classification performance.ResultsExperimental results demonstrate that the proposed MMCN_TCMB model achieves superior performance in multimodal fake news detection. The model attains a recall of 97.988%, precision of 96.223%, F1-score of 97.098%, and overall accuracy of 97.436%, outperforming existing methods.DiscussionThe findings indicate that integrating multimodal feature extraction with cross-attention mechanisms and the TCMB loss function significantly enhances the reliability and accuracy of fake news detection. The proposed framework effectively captures both textual and visual inconsistencies, making it a promising approach for combating misinformation in modern digital platforms.The code is available on:https://github.com/banbhrani84/MMCN_TCMB-Fake-News-.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1799073</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1799073</link>
        <title><![CDATA[The role of statistical methods and artificial intelligence in inventory management for manufacturing industries: a systematic literature review]]></title>
        <pubdate>2026-05-15T00:00:00Z</pubdate>
        <category>Systematic Review</category>
        <author>Arvia Dwi Royani</author><author>Mahfud Sholihin</author><author>Dewi Dewi</author><author>Novika Novika</author><author>Annisa Sorayya</author><author>Wahyu Nur Hanifah</author><author>Rizki Ramadhani Arif Trilana</author><author>Paolina Buton</author>
        <description><![CDATA[Inventory management is a critical business process that affects the operational efficiency and competitiveness of manufacturing companies. Inaccurate inventory decisions can result in significant financial losses for companies. Demand variability poses a challenge in determining inventory levels, requiring more sophisticated, flexible forecasting methods. This study was conducted to examine the roles of statistical methods and Artificial Intelligence (AI) in inventory decision-making in the manufacturing industry, analyze the conditions under which each method is suitable, and evaluate the potential of a hybrid approach integrating statistical methods and AI. This study uses the Systematic Literature Review method with the PRISMA 2020 framework to ensure research transparency and accuracy. This study identifies articles from reputable databases indexed in Scopus. The findings show a significant shift in inventory management research. In the last decade, AI technology has dominated the literature at 62.5%, while statistical methods account for 25%, and hybrid methods have begun to emerge but remain limited to 12.5%. Based on the review of selected papers, statistical methods have proven to remain effective for consistent historical data and stable demand patterns. Conversely, in dynamic operational environments with large-scale data and complex nonlinear patterns, AI technology is superior. This study also found that the hybrid approach has great potential to balance accuracy, interpretability, and decision support, although the relevant literature remains limited. The implementation of technology in the manufacturing industry faces several obstacles, including limited data quality, a skills gap in technology, and the black-box nature of complex AI. This review provides a systematic and critical synthesis of methodological patterns and operational fit in the use of statistical, AI, and hybrid methods for manufacturing inventory management. Future research is recommended to focus on the development of interpretable AI, modular hybrid frameworks, and the use of real industry data to ensure that academic innovations can be applied in the manufacturing industry.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1817120</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1817120</link>
        <title><![CDATA[Cheatomaly: weakly supervised video anomaly ranking for exam cheating detection using vision transformers]]></title>
        <pubdate>2026-05-12T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>El Mehdi Alaoui Mrani</author><author>Anas Bouayad</author><author>Khalid Fardousse</author>
        <description><![CDATA[Detecting cheating in classroom examinations is challenging because suspicious behaviors are often subtle, temporally sparse, and context-dependent. To address the lack of dedicated benchmarks for this setting, we introduce Cheatomaly, a curated video dataset assembled from publicly available classroom examination material and annotated to support weakly supervised anomaly detection. We formulate cheating detection as a weakly supervised video anomaly ranking task using Multiple Instance Learning (MIL) with Vision Transformer features. Videos are divided into temporal segments, and segment-level representations are built using mean pooling and a Mean, Standard Deviation, and Temporal Difference (MSD) formulation. A margin-based ranking objective is used to prioritize anomalous videos and suspicious temporal segments using only video-level labels during training. Experimental results on Cheatomaly show strong video-level discrimination and meaningful frame-level localization across repeated runs. Ablation, baseline, statistical, and sensitivity analyses indicate that temporal aggregation affects the trade-off between ranking and localization but does not produce consistent statistically significant gains. Overall, Cheatomaly provides a realistic benchmark for studying subtle cheating-related anomalies in classroom examinations, and the results highlight that the main challenge lies in modeling context-dependent temporal behavior rather than feature aggregation alone.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1821612</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1821612</link>
        <title><![CDATA[A longitudinal multimodal big data infrastructure for precision poultry monitoring]]></title>
        <pubdate>2026-05-11T00:00:00Z</pubdate>
        <category>Methods</category>
        <author>Daniel Essien</author><author>Yashan Dhaliwal</author><author>Suresh Neethirajan</author>
        <description><![CDATA[Livestock systems are increasingly instrumented with heterogeneous sensors, yet the resulting data remain fragmented, short-lived, and rarely documented as integrated infrastructures. This gap limits the development of robust multimodal artificial intelligence under real production conditions. Here we present a longitudinal multimodal data infrastructure for poultry monitoring, spanning 22 consecutive weeks across five commercial-style barns. The dataset combines continuous RGB video (1080 p, 30 fps), continuous audio (48 kHz), periodic radiometric thermal imaging, and twice-daily environmental measurements, yielding 10.2 terabytes of temporally heterogeneous data. Rather than focusing on a specific predictive task, the study addresses the underlying data-engineering challenge: how to acquire, synchronize, store, and preprocess multimodal streams at production scale. We detail a reproducible system architecture for distributed sensing, local buffering, secure transfer, and cloud-based organization, together with standardized preprocessing pipelines for illumination correction, acoustic denoising, and radiometric temperature extraction. Temporal alignment is achieved through timestamp-based normalization across asynchronous modalities, with explicit characterization of alignment granularity and missing data under real-world constraints. This work positions multimodal livestock sensing as a data-systems problem. The resulting dataset supports longitudinal analysis, cross-modal querying, and the development and evaluation of machine learning and multimodal fusion approaches at appropriate temporal scales. By releasing both data and workflows, we provide a transparent and extensible foundation for building and evaluating AI systems in precision agriculture.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1768571</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1768571</link>
        <title><![CDATA[Optimizing large-scale graph database ingestion through edge value ranking: a proposed framework]]></title>
        <pubdate>2026-05-11T00:00:00Z</pubdate>
        <category>Hypothesis and Theory</category>
        <author>Phanindra Reddy Madduru</author><author>Bijo Thomas</author>
        <description><![CDATA[This paper proposes a preprocessing framework for optimizing large-scale graph database ingestion through intelligent edge filtering based on value ranking. We combine adapted PageRank algorithms with business-specific metrics and edge type importance to evaluate and rank edges, enabling selective retention of high-value relationships. The framework introduces three PageRank variants (maximum weight normalization, weighted average, and log-based normalization) with type-specific business value normalization to handle heterogeneous graphs. Current graph database ingestion approaches struggle with scale: loading 6.2TB of data (38 billion objects) requires over 3 weeks, forcing organizations to limit historical data retention. Our approach addresses this through preprocessing-stage filtering before database ingestion. While requiring experimental validation, preliminary analysis suggests potential for 40%–80% data volume reduction depending on graph characteristics, with corresponding improvements in loading efficiency and storage costs. The paper details the theoretical framework, computational complexity analysis, formal property preservation guarantees, and comprehensive validation methodology. This work represents a novel direction in graph database optimization: value-based preprocessing rather than runtime query optimization.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1807184</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1807184</link>
        <title><![CDATA[Explainable gradient convolutional vector fuzzy pattern analysis based on ensemble model for facial expression recognition]]></title>
        <pubdate>2026-05-08T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Lakshmi Sarvani Videla</author><author>Babu Reddy Mukamalla</author>
        <description><![CDATA[Facial expression recognition using machine learning involves training algorithms to identify and categorize human emotions based on visual cues from facial features. Explainable AI (XAI) enhances this process by providing transparency into how these algorithms arrive at their predictions. While machine learning algorithms provide the capability to recognize facial expressions, explainable AI offers the crucial ability to understand and interpret these recognition processes, leading to more robust, fair, and trustworthy systems.The aim of this research is to propose a novel method in facial expression recognition using segmentation by an ensemble machine learning algorithm and explainable AI model. The input consists of facial expression images, which are first processed for noise removal and normalization. The processed images are then segmented using the Explainable Gradient Convolutional Vector Fuzzy Pattern Recognition (ExGrConVFuzPR) model.The proposed method was evaluated on the JAFFE, CK, and AFLW datasets. The model achieved promising results with an accuracy of 97%, precision of 96%, recall of 96%, F1-score of 97%, and RMSE of 0.043. These outcomes demonstrate that the suggested approach provides good performance along with improved interpretability.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1786859</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1786859</link>
        <title><![CDATA[Detection and classification of lung cancer using sequential hybridization of CNN and RNN type architectures]]></title>
        <pubdate>2026-05-08T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Maheswari Vutukuri</author><author>Parveen Sultana Habibullah</author>
        <description><![CDATA[IntroductionEarly and accurate lung cancer detection from computed tomography (CT) images remains a challenging task because of the complex morphology of lung nodules, class imbalance, variation in image quality, and the risk of overfitting in deep learning models. Conventional manual interpretation is time-consuming and may be affected by inter-observer variability. Therefore, an automated and reliable CT-based classification framework is required to support early identification of benign, malignant, and normal lung conditions.MethodsThis study proposes a sequential hybrid deep learning framework that integrates convolutional and recurrent neural network components for multiclass lung cancer classification. A dataset of 1,600 CT-scan images collected from multiple hospital data repositories across Bengaluru was used and divided in an approximate 70:30 ratio for training and validation. The preprocessing pipeline includes contrast enhancement using Contrast Limited Adaptive Histogram Equalization (CLAHE), morphological operations for lung segmentation, and nodule-focused masking to isolate diagnostically relevant lung regions. Data augmentation and transfer learning were applied to improve model generalization and reduce overfitting. DenseNet201 was used for feature extraction, while a bidirectional gated recurrent unit (BiGRU) module was incorporated for sequential representation learning. Hyperparameter optimization and early stopping were used to improve training stability and classification performance.ResultsThe proposed DenseNet201-BiGRU sequential hybrid architecture achieved an overall accuracy of 95.8%. The class-wise accuracies were 97.33% for benign cases, 93.33% for malignant cases, and 96.67% for normal cases. Precision, recall, and F1-score values further demonstrated that the model maintained reliable classification performance across diverse and imbalanced CT image classes.DiscussionThe results indicate that sequential hybridization of DenseNet201-based feature extraction with BiGRU-based representation learning provides an efficient, precise, and robust framework for CT-based lung cancer classification. The proposed method improves classification reliability by combining enhanced preprocessing, focused lung-region extraction, transfer learning, and recurrent modeling. However, further validation using larger, multi-center datasets and additional clinical testing is required before real-world deployment and broader diagnostic application.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1761377</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1761377</link>
        <title><![CDATA[The designing of a transparent hybrid machine learning framework for water leak detection: a systematic review]]></title>
        <pubdate>2026-04-24T00:00:00Z</pubdate>
        <category>Systematic Review</category>
        <author>Chinemerem M. Anozie</author><author>Tite Tuyikeze</author><author>Ibidun C. Obagbuwa</author><author>Fezile Matsebula</author>
        <description><![CDATA[IntroductionGlobal water scarcity is increasingly exacerbated by substantial water losses, with approximately 30% of treated water lost annually due to leaks in aging Water Distribution Networks (WDNs). Addressing this challenge requires advanced and reliable leak detection mechanisms. This study investigates the design of a transparent hybrid machine learning framework aimed at improving the accuracy and effectiveness of water leak detection systems.MethodsA systematic literature review was conducted following PRISMA guidelines. A total of 27 relevant studies were analyzed, focusing on hybrid deep learning approaches that incorporate data fusion, mixed models, and ensemble techniques for leak detection in WDNs.ResultsThe findings indicate that hybrid and ensemble learning techniques are becoming more important in the identification of water leaks. Several studies reported exceptional high performance, with some models achieving up to 99% balanced accuracy by leveraging multiple data modalities. These approaches demonstrate strong resilience and adaptability across varying operational conditions.DiscussionDespite their high performance, the complexity and “black-box” nature of hybrid models limit their practical deployment. The study highlights the importance of integrating Explainable Artificial Intelligence (XAI) techniques to enhance transparency, interpretability, and user trust. The review concludes that future intelligent leak management systems should combine high-performing hybrid models with XAI to develop efficient, interpretable, and trustworthy decision-support systems that support sustainable water resource management.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1772101</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1772101</link>
        <title><![CDATA[Generation of Kazakhstan's unified national testing variants using AI: a platform for automatic task creation with expert control]]></title>
        <pubdate>2026-04-17T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Bolatbek Abdrasilov</author><author>Talgat Niyazov</author><author>Lyazzat Shinetova</author><author>Shugyla Altybayeva</author><author>Kenzhekul Turalbayeva</author><author>David Orlov</author>
        <description><![CDATA[This study examines the use of artificial intelligence for Automatic Item Generation (AIG) in the context of Kazakhstan's Unified National Testing (UNT) and presents a human-in-the-loop platform for scalable, expert-controlled test development. The objective is to evaluate whether large language models (LLMs) can reliably generate high-quality, isomorphic mathematics test items in the Kazakh language while preserving psychometric and pedagogical requirements. A hybrid AI system combining a local and a cloud-based LLM was implemented to perform semantic deconstruction of prototype items and constrained isomorphic generation of new variants. The pipeline included structured prompt engineering, parallel generation, and automated symbolic validation using Python and SymPy, followed by double-blind expert review. A stratified sample of 120 UNT mathematics items served as prototypes, from which 200 AI-generated clones were produced and validated. Six qualified subject-matter experts conducted independent evaluations using standardized criteria. Inter-rater reliability reached a substantial level (Cohen's κ = 0.78). Results show that 97.5% of generated items were recommended for use after review, with 50.5% accepted without revision and 47.0% accepted after corrections. The most frequent revision needs involved difficulty calibration, wording clarity, and factual or curricular alignment. Expert interviews confirmed that AI generation significantly reduces development time but remains limited in higher-order cognitive item design and pedagogically grounded distractor construction, especially in a low-resource, morphologically complex language environment. The findings support a hybrid augmentation model in which AI accelerates large-scale item production while experts ensure linguistic, cultural, and psychometric validity. The proposed framework demonstrates practical potential for multilingual, high-stakes assessment systems and provides implementation guidelines for responsible AI integration in test development.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1778363</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1778363</link>
        <title><![CDATA[Tree-based machine learning methods for predicting vehicle insurance claim size]]></title>
        <pubdate>2026-03-23T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Edossa Merga Terefe</author><author>Merga Abdissa Aga</author>
        <description><![CDATA[Vehicle insurance claim severity modeling requires accurate and interpretable methods that can handle skewed and heterogeneous loss data. This study provides a structured empirical comparison between classical parametric regression models and tree-based ensemble learning approaches for predicting claim size conditional on claim occurrence. The analysis is conducted within a cross-sectional conditional severity framework using real-world motor insurance data. We implement and compare ordinary least squares (OLS), a Tweedie generalized linear model (GLM), and three ensemble methods: bagging, random forests (RFs), and gradient boosting. Model performance is evaluated using out-of-sample root mean square error (RMSE), and variable importance measures assess the relative contribution of predictors. The results indicate that tree-based ensemble methods achieve modest improvements in predictive accuracy relative to classical parametric models. The Tweedie GLM remains a competitive, flexible parametric benchmark for skewed positive claim amounts. Variable importance analysis consistently identifies premium and insured value as key determinants of claim severity. Overall, the findings suggest that ensemble learning methods can complement traditional actuarial models, offering additional flexibility in capturing non-linear effects while maintaining comparable predictive performance in moderate-complexity severity data.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1676922</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1676922</link>
        <title><![CDATA[FunduScope: a human-centered, machine learning–based interactive tool for training junior ophthalmologists in diabetic retinopathy detection]]></title>
        <pubdate>2026-03-13T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Sara-Jane Bittner</author><author>Michael Barz</author><author>Daniel Sonntag</author>
        <description><![CDATA[Interpreting fundus images is an essential skill for detecting eye diseases, such as diabetic retinopathy (DR), one of the leading causes of visual impairment. However, the training of junior doctors relies on experienced ophthalmologists, who often lack the time for teaching, or on printed training materials that lack variability in examples. In this work, we present FunduScope, an interactive human-centered learning tool for training junior ophthalmologists, which is based on a pre-trained ML model for classifying DR. In a qualitative pre-study, we investigated the needs of junior doctors and identified gaps in recent learning procedures. In the main mixed-methods study, we examined the experience of 10 junior doctors with the tool and its impact on cognitive load, usability, and additional factors relevant to e-learning tools. Despite technical constraints our results confirm the potential of using an ML-based learning tool in medical education, addressing the time constraints of ophthalmologists, and providing learning independence for junior doctors. However, future work could extend the learning tool by using explainable artificial intelligence (XAI) to further support the clinical decision making of learners and exceeding the scope of this proof of concept to other ophthalmic diseases.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1718710</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1718710</link>
        <title><![CDATA[Modeling household adoption of IoT-based home security in Dhaka: a PLS–machine learning framework]]></title>
        <pubdate>2026-02-04T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Arif Mahmud</author><author>Ashikur Rahman</author><author>Fahmid Al Farid</author><author>Jia Uddin</author><author>Hezerul Bin Abdul Karim</author>
        <description><![CDATA[IntroductionDespite several strategies, Bangladesh has a poor rate of internet of things (IoT) deployment. This study therefore seeks to investigate the factors shaping IoT adoption for residential security in Dhaka and to analyze their respective contributions.MethodHence, this study combined two important theories, namely protection motivation theory (PMT) along with attitude-social influence-self-efficacy (ASE) in which a hybrid PLS-Machine learning approach has been used to identify both linear and nonlinear correlations with high predictive accuracy. Snowball sampling method was utilized to choose 348 valid replies from a survey of household heads. Afterward, partial least squares (PLS) followed by artificial neural networks (ANN) and machine learning (ML) classifiers were the procedures that made up the complete assessment method.ResultsThe variables that affected intention with a variance of 34.9% and accuracy of 74.28% were severity, vulnerability, response efficacy, response cost, and attitude. On the other hand, vulnerability was the most significant predictor, followed by response cost, attitude, response efficacy, self-efficacy, social influence, and severity.DiscussionThe theoretical contribution of this study lies in its novel integration of PMT and ASE models, offering new insights into their combined effect on technology adoption in emerging markets. Besides, the findings contribute to the literature by increasing the public awareness of home security that can enhance Dhaka's overall state of public order and safety. Moreover, the findings may offer valuable insights for companies and entrepreneurs, as incorporating these factors into marketing strategies and investment initiatives is likely to foster greater consumer adoption.]]></description>
      </item>
      </channel>
    </rss>