<?xml version="1.0" encoding="utf-8"?>
    <rss version="2.0">
      <channel xmlns:content="http://purl.org/rss/1.0/modules/content/">
        <title>Frontiers in Big Data | Machine Learning section | New and Recent Articles</title>
        <link>https://www.frontiersin.org/journals/big-data/sections/machine-learning</link>
        <description>RSS Feed for Machine Learning section in the Frontiers in Big Data journal | New and Recent Articles</description>
        <language>en-us</language>
        <generator>Frontiers Feed Generator,version:1</generator>
        <pubDate>2026-09-13T00:12:43.33+00:00</pubDate>
        <ttl>60</ttl>
        <item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1917723</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1917723</link>
        <title><![CDATA[Data-driven adaptive hybrid models for exchange rate return forecasting]]></title>
        <pubdate>2026-09-11T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Olumide Sunday Adesina</author><author>Lawrence Ogechukwu Obokoh</author>
        <description><![CDATA[BackgroundExchange-rate return forecasting is challenging because financial time series may exhibit linearity, nonlinearity, regime-switching behavior, and volatility. To address these complexities, two adaptive hybrid forecasting frameworks were developed: ATW-HyF A, which dynamically combines ARIMA, SETAR, and ANN forecasts using inverse-variance weighting, and ATW-HyF B, which extends the framework by incorporating a GARCH(1,1) volatility layer to model the conditional variance of forecast errors.MethodsThe robustness of the proposed frameworks was assessed using simulated data and three foreign exchange return series: USD/NGN, EUR/USD, and GBP/USD. Forecast performance was evaluated within a rolling-origin forecasting framework and compared with competing forecasting models.ResultsThe simulation results showed that the ATW-HyF models achieved the lowest forecast errors among the competing models. However, their performance on the empirical exchange-rate series was mixed, with other models achieving lower forecast errors for some currency pairs. In particular, the ARIMA-ANN hybrid performed best for USD/NGN and EUR/USD, whereas ATW-HyF B performed best for GBP/USD.DiscussionThe findings indicate that forecasting performance depends on the characteristics of the underlying exchange-rate series rather than on model complexity alone. Combining models with complementary predictive strengths can improve forecasting performance, particularly when the data exhibit linear, nonlinear, and volatility-related patterns. The incorporation of adaptive weighting and volatility modeling therefore provides a flexible approach to exchange-rate return forecasting under evolving market conditions.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1900836</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1900836</link>
        <title><![CDATA[Fine-tuning multilingual sentence transformers for low-resource skill-concept matching across Kazakh, Russian, and English]]></title>
        <pubdate>2026-09-10T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Sandugash Serikbayeva</author><author>Madina Sambetbayeva</author><author>Valiya Ramazanova</author><author>Aigerim Yerimbetova</author><author>Zhanar Lamasheva</author><author>Zhanna Sadirmekova</author><author>Ardak Batyrkhanov</author><author>Yersaiyn Mailybayev</author>
        <description><![CDATA[IntroductionMatching short, specialized skill expressions across English, Russian, and Kazakh is challenging because general-domain multilingual encoders underperform on terse, code-mixed, domain-specific phrases, particularly in the low-resource Kazakh setting.MethodsWe fine-tuned a multilingual Sentence Transformer using a staged Multiple Negatives Ranking objective on a trilingual paraphrase corpus, including Russian augmentation pairs. We evaluated skill matching and semantic similarity across languages and assessed the resulting embeddings through downstream skill-taxonomy clustering.ResultsFine-tuning preserved English performance while improving Russian and Kazakh similarity quality. Russian cosine Pearson correlation increased from 0.8125 to 0.8221, while Kazakh cosine Pearson increased from 0.5989 to 0.6050 and Kazakh dot-product Pearson from 0.4487 to 0.4912. Agglomerative clustering improved mean silhouette from 0.27126 to 0.2851 and reduced erroneous clusters from 19.74% to 13.97%.DiscussionThe results provide evidence consistent with cross-lingual transfer as an important mechanism of improvement for Kazakh. They also motivate language-specific threshold calibration and demonstrate that intrinsic similarity improvements translate into a cleaner downstream skill taxonomy.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1876811</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1876811</link>
        <title><![CDATA[Reconstruction-aware urban crime analytics from incomplete judicial records using heterogeneous graph temporal imputation]]></title>
        <pubdate>2026-09-07T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Wen Xu</author><author>Yong Dai</author><author>Qi Zhang</author><author>Keyu Chen</author><author>Sanxia Zeng</author>
        <description><![CDATA[IntroductionIncomplete judicial records constrain reliable crime analytics because missing attributes, heterogeneous case descriptions, mixed variable types, and imbalanced charge distributions reduce their analytical value. This study developed a reconstruction-aware framework for drug-crime charge classification from incomplete judicial records.MethodsDrug-related judicial documents from Chengdu, China, covering 2014-2021 were screened, cleaned, and converted into structured person-level observations. The final dataset comprised 12,620 valid documents and 15,184 observations with 16 structured features. A Heterogeneous Graph Convolution Temporal Autoencoder (HetGConv-TAE) was developed to jointly model heterogeneous case-attribute relations and temporal changes in case composition. Performance was evaluated under 10%, 20%, and 30% controlled missingness, followed by XGBoost charge classification and SHAP interpretation.ResultsHetGConv-TAE achieved the best mean reconstruction performance across conventional, static graph, temporal, and relational graph baselines. XGBoost trained on HetGConv-TAE reconstructed data achieved weighted F1 scores close to 80% and ROC-AUC values above 90%. SHAP analysis identified drug weight, place category, and administrative district as the most influential predictors across charge categories.DiscussionCombining label-excluded heterogeneous graph reconstruction, temporal encoding, downstream classification, and interpretable analysis improves the analytical usefulness of incomplete judicial archives. The framework is intended to support data-quality assessment and aggregate public-safety research rather than automated legal decision-making.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1888885</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1888885</link>
        <title><![CDATA[Study of deep learning cues for cross linguistic part of speech tagging in English– Malayalam code-mixed data]]></title>
        <pubdate>2026-08-28T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Parvathy Padmakumar</author><author>Shreya S. Nair</author><author>A. K. Prajisha</author><author>S. Thara</author>
        <description><![CDATA[IntroductionPart Of Speech (POS) tagging is a fundamental task in Natural Language Processing (NLP) that assigns grammatical labels to words in a sentence. Code mixed text, which entails switching between two or more languages within a single conversation or a sentence, presents challenges for POS tagging. This investigation entailed a comprehensive study of deep learning approaches for cross linguistic POS tagging, focused on English Malayalam code mixed data prevalent on social media platforms. The study was carried out on linguistically complex and varied English Malayalam code mixed text from social media platforms with informal spellings, language switching, transliteration, slang, and unclear grammatical boundaries, reflecting the characteristics of informal online communication. We provide the first structured analysis of micro macro divergence and rare class behavior in English Malayalam code mixed POS tagging.MethodsWe evaluated 14 state of the art model configurations that span traditional sequence labeling approaches and multilingual transformer architectures. Models were compared using standard performance metrics prevalent in the domain of data science, supplemented by normalized confusion matrices, error prone tag identification and micro-macro F1 gap analysis.ResultsOur results showed that CRF (No Lang) emerged as the most balanced model overall on macro F1 (all classes) of 0.8170, while (BiLSTM + CRF) achieved the highest macro F1 (seen classes) of 0.8831, precision of 0.9167, and recall of 0.875, though this reflects strong performance concentrated on frequent tag classes rather than balanced coverage across the full tag set. Notably, the pretrained multilingual transformers (mBERT, MuRIL), despite prior exposure to Malayalam during pretraining, were outperformed on several key metrics by CRF and BiLSTM models trained directly on the code mixed dataset.DiscussionThis finding was contrary to our expectation that existing multilingual knowledge would translate into a clear advantage on this task and merits further investigation.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1889044</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1889044</link>
        <title><![CDATA[A cognitively inspired feature-level fusion framework for interpretable retail sales forecasting using integration of extreme gradient boost machine, artificial neural network and attention mechanism model]]></title>
        <pubdate>2026-08-20T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Munienge Mbodila</author><author>Omobayo Ayokunle Esan</author>
        <description><![CDATA[IntroductionAccurate retail sales forecasting in modern data-intensive environments requires models that not only achieve high predictive accuracy but also scale efficiently as data volume and feature complexity increase. This study proposes a novel cognitively inspired hybrid framework, XGB–ANN–Attn, for interpretable and scalable retail analytics.MethodsThe model introduces feature-level fusion by integrating XGBoost-derived leaf embeddings with deep neural representations, enabling joint modeling of low-order statistical dependencies and high-order nonlinear feature interactions. A lightweight attention mechanism dynamically assigns instance-specific feature importance, enhancing both predictive performance and interpretability while maintaining linear computational complexity with respect to feature dimensionality. Unlike conventional ensemble approaches that operate at the decision level, the proposed framework integrates representations at the level of representations, reducing redundancy and improving computational efficiency. The model is designed for scalability, combining the log-linear complexity of gradient boosting with the linear scaling properties of neural networks and attention mechanisms.ResultsExperimental evaluation on the BigMart and Walmart datasets demonstrates superior performance compared to state-of-the-art models, achieving RMSE = 0.1584 and R2 = 0.9946 on BigMart, and RMSE = 0.8652 and R2 = 0.9568 on Walmart. Furthermore, the framework supports parallelization and distributed deployment, making it suitable for large-scale retail systems.DiscussionThe alignment between attention weights and SHAP explanations provides transparent and actionable insights. The results confirm that the proposed approach offers a scalable, interpretable, and high-performance solution for AI-driven decision support systems in big-data retail environments.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1884673</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1884673</link>
        <title><![CDATA[Benchmarking retrieval augmented generation LLMs for Arabic noise robustness]]></title>
        <pubdate>2026-08-17T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Saleh Almohaimeed</author><author>Abdulrahman Alabduljabbar</author><author>Mousa Jari</author><author>Mohammed Alkhowaiter</author><author>Saad Almohaimeed</author><author>Mohamad Mahmoud Al Rahhal</author>
        <description><![CDATA[Hallucination has become a serious concern in large language models (LLMs), as these models can generate useful yet incorrect or misleading information, which has led to growing research interest in retrieval-augmented generation (RAG) as a mitigation approach. RAG provides LLMs with access to external information, such as databases or documents, which can help them to answer users' questions. Currently, several benchmarks released measure RAG performance on various LLMs; however, evaluations of the noise robustness ability in Arabic are absent. In this paper, we systematically investigated the capabilities of state-of-the-art multilingual LLMs with regard to two essential RAG abilities, noise robustness and negative rejection. To accomplish this, we generated an Arabic benchmark consisting of 300 questions along with 6,196 documents. Then, we assessed the performance of six LLMs in relation to the two aforementioned RAG abilities. The results reveal that all six LLMs were negatively affected when the noise ratio in the external documents was increased. Under the highest noise settings at 80%, the best LLM performance was for Claude-4 sonnet, in which their performance decreased by only 4.67 percentage points. Furthermore, when it comes to the negative rejection task, there has been a significant impact on all six models. The best two models, Claude-4 sonnet and Llama-4, scored 90.67% and 85.67%, respectively, while smaller models, like GPT-3.5, scored 69.33%. Furthermore, our manual analysis reveals that many errors made by LLMs are attributed to over-caution behavior. LLMs often decline to respond probably due to training mechanisms designed to reduce hallucinations. Additionally, other errors occur when there is a high lexical similarity between the question and the words of noisy documents, which causes the model to rely on irrelevant content instead of the correct information.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1916523</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1916523</link>
        <title><![CDATA[Differentiated memory and scientific cognition in AI research agents]]></title>
        <pubdate>2026-08-05T00:00:00Z</pubdate>
        <category>Hypothesis and Theory</category>
        <author>Diego F. Cuadros</author><author>Abdoul-Aziz Maiga</author><author>Sid Thatham</author><author>Alvaro Ortiz</author><author>Margaret Powers-Fletcher</author><author>Ming Tang</author>
        <description><![CDATA[AI research agents increasingly support ideation, literature search, coding, experimental execution, analysis, and manuscript drafting across the scientific workflow. This progress advances automated discovery, but workflow automation is not scientific cognition. Scientific reasoning is cumulative and path-dependent: it depends on what a researcher has written, read, learned from critique, absorbed through experience, and used as habitual standards for judging novelty, rigor, feasibility, and significance. We propose Mnemo as a framework for modeling scientific cognition in AI research agents. First, scientific cognition may require differentiated memory, organized into distinct spaces for authored work, external reference, critique, experience, and judgment. Second, provenance should be treated not as passive metadata but as memory routing, because source origin helps determine cognitive function in reasoning. Third, new ideas may be better modeled as controlled collisions across memory spaces, filtered by judgment, rejection, and epistemic calibration, than as generic recombination from model priors. Mnemo motivates a research agenda for AI in science centered on routing fidelity, critique use, judgment alignment, rejection quality, and epistemic calibration.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1832790</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1832790</link>
        <title><![CDATA[Machine learning approach for predicting the severity risk of obstructive sleep apnea syndrome]]></title>
        <pubdate>2026-07-16T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Qi Wang</author><author>Xiaoyu Yang</author><author>Shuran Xu</author><author>Haohao Wu</author><author>Guixuan Wang</author><author>Huixian Liu</author><author>Ronghua Chen</author><author>Fengming Xu</author><author>Cheng Wang</author><author>Kang Du</author>
        <description><![CDATA[BackgroundObstructive Sleep Apnea-Hypopnea Syndrome (OSAHS) has a high global prevalence and is prone to causing various serious complications. Our objective is to develop severity stratification of OSAHS by integrating multiple commonly available clinical features based on machine learning (ML).Materials and methodsThis study collected data from 432 cases at Qujing Central Hospital in Yunnan Province, integrating 25 clinical feature variables. The cases were randomly split into training (70%) and validation (30%) sets. The importance of the 25 features was analyzed.ResultsIt showed that the HCY, TBIL, BMI, GGT, and Age made significant contributions to OSAHS severity. We established five machine learning models-Multilayer Perceptron (MLP), Random Forest, XGBoost, LightGBM, and Support Vector Machine (SVM)-by integrating 25 clinical features. Through cross-validation and continuous adjustment of model parameters, the optimal predictive model was determined. By calculating model accuracy and F1-score, XGBoost was identified as the best-performing model, achieving an area under the curve (AUC) of 0.63, an accuracy of 75% and an F1-score of 65.60.ConclusionIn this study, we established a predictive model for the severity stratification of OSAHS based on machine learning algorithms. The XGBoost model demonstrated superior predictive performance.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1837706</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1837706</link>
        <title><![CDATA[Noise-robust temporal–spectral fusion transformers for EEG-based cognitive state classification in aviation environments]]></title>
        <pubdate>2026-07-09T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Quynh Anh Nguyen</author><author>Nam Anh Dao</author><author>Long Nguyen</author>
        <description><![CDATA[Attention-related Pilot Performance Decrements (APPD) contribute substantially to aviation incidents, yet existing electroencephalography (EEG)-based monitoring methods often lack generalization, robustness to noise, and effective temporal–spectral integration. We propose a temporal–spectral fusion transformer (TF-T) combining multi-scale preprocessing, dual-stream temporal and spectral feature extraction, and transformer-based fusion with enhanced temporal–spectral integration and multi-resolution feature processing for multiclass cognitive state recognition. Three variants (TF-T1–TF-T3) are evaluated on controlled and ecologically realistic EEG datasets under clean and noise-augmented (Gaussian, Uniform, COMBO) conditions, using chronological partitioning to avoid temporal leakage. TF-T2 achieves the highest clean-data accuracy (99.2%), while TF-T3 offers superior robustness, improving Macro-F1 by ~4.5–4.7 points across all noise types and outperforming state-of-the-art baselines by up to +8 Macro-F1 under COMBO noise, supporting its deployment in perturbation-prone aviation environments.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1842233</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1842233</link>
        <title><![CDATA[Evolutionary multi-agent reinforcement learning for crisis-aware demographic policy optimization]]></title>
        <pubdate>2026-07-08T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Anton V. Dozhdikov</author><author>Arseniy M. Sitkovskiy</author>
        <description><![CDATA[Demographic systems face unprecedented challenges from simultaneous crises. Conventional statistical demography techniques and agent–based models often struggle to capture nonlinear inter–regional interactions during periods of severe socio–economic disruption. To address this, we propose MADDPG–EVO–DGM, a hybrid algorithm that integrates multi–agent deep reinforcement learning with evolutionary optimisation and meta–learning principles to model regional demographic processes under multiple crisis scenarios. Each region is treated as an autonomous agent learning to steer demographic policy levers, while periodic evolutionary “boosters” overcome local optima via population–based perturbations of actor network parameters. Additionally, a Darwin–Gödel Machine–inspired meta–learning mechanism adapts the booster triggers, enabling self–improvement in the learning process. We evaluate MADDPG–EVO–DGM on a simulation environment calibrated with real demographic data for eight federal regions of the Russian Federation over the period 2000–2024 and subject to ten concurrent crisis scenarios (e.g., pandemic, geopolitical conflict, economic collapse). Experiments demonstrate significantly faster convergence and improved performance over a baseline MADDPG: the hybrid approach achieves a higher final average reward (252.57 vs. 243.07) and 3.4 × lower convergence variance (σ = 0.24 vs. 0.80), indicating more reliable training. It also exhibits qualitative performance jumps of +68% during evolutionary phases and maintains 35%–45% greater resilience under crisis shocks compared to the baseline. To our knowledge, this is the first application of multi–agent reinforcement learning to large–scale demographic modeling under crises, opening new possibilities for evidence–based, crisis–resilient population policy design. Code, data, and logs are provided to ensure reproducibility.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1838191</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1838191</link>
        <title><![CDATA[Measuring the impact of virtualization and containerization on the environment when using GPUs for processing the AI models]]></title>
        <pubdate>2026-06-16T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Safaa Hriez</author><author>Mohammad Haikal</author>
        <description><![CDATA[IntroductionThe rapid growth of artificial intelligence (AI) has significantly increased computing demand, intensifying the operational strain on the computing environment. While virtualization and containerization are established technologies for resource optimization, their comparative energy efficiency and environmental impact, particularly under GPU-accelerated AI workloads, are not well-quantified.MethodsThis study evaluates the energy consumption and environmental impact of virtualization and containerization technologies when using Graphics Processing Units (GPUs) for AI model execution. Employing a computer vision benchmark, the performance, GPU resource utilization, and power consumption were measured. The experiment involved training a DenseNet-121 model on the MNIST dataset within a VirtualBox virtual machine and a Docker container environment.ResultsThe analysis indicates that containerization consistently surpasses virtualization in energy efficiency. Specifically, Docker container configuration demonstrated an approximately 21.6% reduction in total energy consumption, and a corresponding reduction in carbon dioxide (CO2) emissions compared to a VirtualBox virtual machine. Furthermore, containerization exhibited lower average and peak GPU utilization and power consumption.DiscussionThese findings demonstrate that containerization offers a more energy-efficient and environmentally sustainable approach than VirtualBox virtualization for the specific GPU-enabled AI workload evaluated in this study. Statistical significance testing indicates that the observed performance differentials are significant, supporting the validity of the results within the experimental scope of this work.ConclusionImplementing containerization in this experimental setup may reduce energy consumption and environmental impact without compromising computational performance. Future studies should extend these analyses to larger neural network models, diverse AI workloads, and heterogeneous GPU platforms to enhance the generalizability of these findings beyond the current single-system experimental configuration.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1752468</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1752468</link>
        <title><![CDATA[Data field theory: a geometric framework for learning on Riemannian manifolds with synthetic validation and limitation analysis]]></title>
        <pubdate>2026-06-11T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Mohammadreza Nehzati</author>
        <description><![CDATA[IntroductionConventional machine learning treats learning as parameter optimization, lacking a first-principles framework for phenomena like criticality, generalization, and causal structure. We introduce Data Field Theory (DFT), a mathematical framework modelling learning as the evolution of a data field governed by stochastic partial differential equations on Riemannian manifolds. This work aims to validate DFT's core predictions in settings where its geometric assumptions hold, while honestly assessing its empirical limitations.MethodsWe formulate learning as a field φ:M×ℝ≥0→ℝk evolving on a spherical manifold. To test DFT, we implement a hierarchical classification task using synthetic data drawn from von Mises-Fisher distributions, ensuring match with the manifold geometry. We derive four key predictions: (1) critical exponents near concept formation, (2) a spectral robustness law linking Eigen gaps to out-of-distribution (OOD) error, (3) finite-speed causal propagation from hyperbolic regularization, and (4) approximate rotational equivariance via a Ward identity. We also conduct a preliminary real-data experiment projecting MNIST digits onto the sphere.ResultsSynthetic experiments validate all four predictions: (1) Correlation length diverges as ξ(t)~|t-tc|-ν with ν = 0.63 ± 0.04, accompanied by 1/f fluctuations; (2) OOD generalization error scales as ϵOOD∝mgap-2 (ρ = −0.78, p < 10−6); (3) Causal propagation speed ceff = 0.98 ± 0.03 (theory maximum cmax = 1.0) under hyperbolic regularization; and (4) Ward identity residual R = 0.0032 ± 0.0008 converging as R∝h1.02. However, on real-world MNIST-sphere data, DFT achieves only 15.7% accuracy versus 51.7% for k-NN, revealing critical limitations.DiscussionDFT successfully predicts emergent phenomena criticality, spectral robustness, bounded causality, and approximate equivariance under ideal geometric conditions, supporting its theoretical validity. The poor real-data performance highlights key gaps: the current framework lacks adaptive metric learning, noise robustness, and hierarchical feature extraction present in real images. These results establish DFT as a principled mathematical foundation for learning as field dynamics while clearly delineating necessary extensions for practical applicability.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1883246</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1883246</link>
        <title><![CDATA[Correction: Explainable gradient convolutional vector fuzzy pattern analysis based on ensemble model for facial expression recognition]]></title>
        <pubdate>2026-06-10T00:00:00Z</pubdate>
        <category>Correction</category>
        <author>Lakshmi Sarvani Videla</author><author>Babu Reddy Mukamalla</author>
        <description></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1825213</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1825213</link>
        <title><![CDATA[When uncertainty guides learning: a highly effective approach to kidney disease classification in CT imaging]]></title>
        <pubdate>2026-06-09T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Muslima Akter</author><author>Fahmid Al Farid</author><author>Md Yousuf Ahmad</author><author>Md Azad Hossain Raju</author><author>Sowad Rahman</author><author>Jia Uddin</author><author>Hezerul Bin Abdul Karim</author>
        <description><![CDATA[The high cost of expert annotations significantly hinders the advancement of deep learning models for clinical medical imaging. This work introduces an efficient entropy-based active learning framework that achieves outstanding classification performance for renal abnormalities (Normal, Cyst, Stone, Tumor) in CT scans while requiring only a minimal amount of labeled data. The dataset comprises 12,446 CT slices split 70/15/15 into training (8,716), validation (1,865), and test (1,865) partitions via stratified sampling. Starting with only 200 randomly selected images and employing predictive entropy for uncertainty sampling on a pretrained ResNet-50 backbone, the proposed method attains 99.71% ± 0.25% mean test accuracy (95% CI: [99.30, 99.94]) across five independent runs after just six query cycles on the standard 12,446-image CT kidney dataset. Our method uses only 2,000 labeled training images, representing 22.9% of the 8,716-image training partition (a 77.1% reduction in required annotations relative to full supervision of the training set). This performance matches or exceeds prior fully supervised methods trained on the complete labeled training partition while demonstrating substantially improved sample efficiency, particularly in early annotation cycles where entropy-guided selection converges significantly faster than random sampling. Statistical testing across five repeated runs confirms that results are stable (Shapiro-Wilk p = 0.148). The framework exhibits exceptional sample efficiency as described by an empirically fitted power-law curve with a fitted exponent of 1.2, and empirically observed uncertainty decay with a rate of 0.92. These results offer both practical insights into annotation efficiency and substantial application value in the medical imaging domain.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1796969</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1796969</link>
        <title><![CDATA[TCMB: cross-model multi-level cross-attention network with Taylor-based loss for multimodal fake news detection]]></title>
        <pubdate>2026-06-05T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Santosh Kumar Banbhrani</author>
        <description><![CDATA[IntroductionThe rapid spread of misinformation across social media platforms, websites, and online communication channels has made fake news detection a critical task in the digital era. Although various computational approaches have been developed to identify fake news, many existing methods suffer from limitations such as biased training datasets and high rates of false positives and false negatives. To address these challenges, this study proposes a Multimodal Cross Attention Network with Taylor-based Cross Entropy Mean Bias (MMCN_TCMB) model for detecting multimodal fake news.MethodsThe proposed approach utilizes multimodal inputs consisting of textual and visual content obtained from fake news datasets. The textual information in news posts is first tokenized using Bidirectional Encoder Representations from Transformers (BERT). Feature extraction is then performed using Word2Vec and Term Frequency–Inverse Gravity Moment (TF-IGM). Simultaneously, images associated with news posts undergo preprocessing through Contrast Limited Adaptive Histogram Equalization and Histogram Equalization (CLAHE-HE), followed by feature extraction using ResNet. The extracted textual and visual features are combined and processed through the MMCN framework. The learning mechanism of the network is enhanced using the Taylor-based Cross Entropy Mean Bias (TCMB) loss function to improve classification performance.ResultsExperimental results demonstrate that the proposed MMCN_TCMB model achieves superior performance in multimodal fake news detection. The model attains a recall of 97.988%, precision of 96.223%, F1-score of 97.098%, and overall accuracy of 97.436%, outperforming existing methods.DiscussionThe findings indicate that integrating multimodal feature extraction with cross-attention mechanisms and the TCMB loss function significantly enhances the reliability and accuracy of fake news detection. The proposed framework effectively captures both textual and visual inconsistencies, making it a promising approach for combating misinformation in modern digital platforms.The code is available on:https://github.com/banbhrani84/MMCN_TCMB-Fake-News-.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1799073</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1799073</link>
        <title><![CDATA[The role of statistical methods and artificial intelligence in inventory management for manufacturing industries: a systematic literature review]]></title>
        <pubdate>2026-05-15T00:00:00Z</pubdate>
        <category>Systematic Review</category>
        <author>Arvia Dwi Royani</author><author>Mahfud Sholihin</author><author>Dewi Dewi</author><author>Novika Novika</author><author>Annisa Sorayya</author><author>Wahyu Nur Hanifah</author><author>Rizki Ramadhani Arif Trilana</author><author>Paolina Buton</author>
        <description><![CDATA[Inventory management is a critical business process that affects the operational efficiency and competitiveness of manufacturing companies. Inaccurate inventory decisions can result in significant financial losses for companies. Demand variability poses a challenge in determining inventory levels, requiring more sophisticated, flexible forecasting methods. This study was conducted to examine the roles of statistical methods and Artificial Intelligence (AI) in inventory decision-making in the manufacturing industry, analyze the conditions under which each method is suitable, and evaluate the potential of a hybrid approach integrating statistical methods and AI. This study uses the Systematic Literature Review method with the PRISMA 2020 framework to ensure research transparency and accuracy. This study identifies articles from reputable databases indexed in Scopus. The findings show a significant shift in inventory management research. In the last decade, AI technology has dominated the literature at 62.5%, while statistical methods account for 25%, and hybrid methods have begun to emerge but remain limited to 12.5%. Based on the review of selected papers, statistical methods have proven to remain effective for consistent historical data and stable demand patterns. Conversely, in dynamic operational environments with large-scale data and complex nonlinear patterns, AI technology is superior. This study also found that the hybrid approach has great potential to balance accuracy, interpretability, and decision support, although the relevant literature remains limited. The implementation of technology in the manufacturing industry faces several obstacles, including limited data quality, a skills gap in technology, and the black-box nature of complex AI. This review provides a systematic and critical synthesis of methodological patterns and operational fit in the use of statistical, AI, and hybrid methods for manufacturing inventory management. Future research is recommended to focus on the development of interpretable AI, modular hybrid frameworks, and the use of real industry data to ensure that academic innovations can be applied in the manufacturing industry.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1817120</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1817120</link>
        <title><![CDATA[Cheatomaly: weakly supervised video anomaly ranking for exam cheating detection using vision transformers]]></title>
        <pubdate>2026-05-12T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>El Mehdi Alaoui Mrani</author><author>Anas Bouayad</author><author>Khalid Fardousse</author>
        <description><![CDATA[Detecting cheating in classroom examinations is challenging because suspicious behaviors are often subtle, temporally sparse, and context-dependent. To address the lack of dedicated benchmarks for this setting, we introduce Cheatomaly, a curated video dataset assembled from publicly available classroom examination material and annotated to support weakly supervised anomaly detection. We formulate cheating detection as a weakly supervised video anomaly ranking task using Multiple Instance Learning (MIL) with Vision Transformer features. Videos are divided into temporal segments, and segment-level representations are built using mean pooling and a Mean, Standard Deviation, and Temporal Difference (MSD) formulation. A margin-based ranking objective is used to prioritize anomalous videos and suspicious temporal segments using only video-level labels during training. Experimental results on Cheatomaly show strong video-level discrimination and meaningful frame-level localization across repeated runs. Ablation, baseline, statistical, and sensitivity analyses indicate that temporal aggregation affects the trade-off between ranking and localization but does not produce consistent statistically significant gains. Overall, Cheatomaly provides a realistic benchmark for studying subtle cheating-related anomalies in classroom examinations, and the results highlight that the main challenge lies in modeling context-dependent temporal behavior rather than feature aggregation alone.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1821612</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1821612</link>
        <title><![CDATA[A longitudinal multimodal big data infrastructure for precision poultry monitoring]]></title>
        <pubdate>2026-05-11T00:00:00Z</pubdate>
        <category>Methods</category>
        <author>Daniel Essien</author><author>Yashan Dhaliwal</author><author>Suresh Neethirajan</author>
        <description><![CDATA[Livestock systems are increasingly instrumented with heterogeneous sensors, yet the resulting data remain fragmented, short-lived, and rarely documented as integrated infrastructures. This gap limits the development of robust multimodal artificial intelligence under real production conditions. Here we present a longitudinal multimodal data infrastructure for poultry monitoring, spanning 22 consecutive weeks across five commercial-style barns. The dataset combines continuous RGB video (1080 p, 30 fps), continuous audio (48 kHz), periodic radiometric thermal imaging, and twice-daily environmental measurements, yielding 10.2 terabytes of temporally heterogeneous data. Rather than focusing on a specific predictive task, the study addresses the underlying data-engineering challenge: how to acquire, synchronize, store, and preprocess multimodal streams at production scale. We detail a reproducible system architecture for distributed sensing, local buffering, secure transfer, and cloud-based organization, together with standardized preprocessing pipelines for illumination correction, acoustic denoising, and radiometric temperature extraction. Temporal alignment is achieved through timestamp-based normalization across asynchronous modalities, with explicit characterization of alignment granularity and missing data under real-world constraints. This work positions multimodal livestock sensing as a data-systems problem. The resulting dataset supports longitudinal analysis, cross-modal querying, and the development and evaluation of machine learning and multimodal fusion approaches at appropriate temporal scales. By releasing both data and workflows, we provide a transparent and extensible foundation for building and evaluating AI systems in precision agriculture.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1768571</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1768571</link>
        <title><![CDATA[Optimizing large-scale graph database ingestion through edge value ranking: a proposed framework]]></title>
        <pubdate>2026-05-11T00:00:00Z</pubdate>
        <category>Hypothesis and Theory</category>
        <author>Phanindra Reddy Madduru</author><author>Bijo Thomas</author>
        <description><![CDATA[This paper proposes a preprocessing framework for optimizing large-scale graph database ingestion through intelligent edge filtering based on value ranking. We combine adapted PageRank algorithms with business-specific metrics and edge type importance to evaluate and rank edges, enabling selective retention of high-value relationships. The framework introduces three PageRank variants (maximum weight normalization, weighted average, and log-based normalization) with type-specific business value normalization to handle heterogeneous graphs. Current graph database ingestion approaches struggle with scale: loading 6.2TB of data (38 billion objects) requires over 3 weeks, forcing organizations to limit historical data retention. Our approach addresses this through preprocessing-stage filtering before database ingestion. While requiring experimental validation, preliminary analysis suggests potential for 40%–80% data volume reduction depending on graph characteristics, with corresponding improvements in loading efficiency and storage costs. The paper details the theoretical framework, computational complexity analysis, formal property preservation guarantees, and comprehensive validation methodology. This work represents a novel direction in graph database optimization: value-based preprocessing rather than runtime query optimization.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1786859</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1786859</link>
        <title><![CDATA[Detection and classification of lung cancer using sequential hybridization of CNN and RNN type architectures]]></title>
        <pubdate>2026-05-08T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Maheswari Vutukuri</author><author>Parveen Sultana Habibullah</author>
        <description><![CDATA[IntroductionEarly and accurate lung cancer detection from computed tomography (CT) images remains a challenging task because of the complex morphology of lung nodules, class imbalance, variation in image quality, and the risk of overfitting in deep learning models. Conventional manual interpretation is time-consuming and may be affected by inter-observer variability. Therefore, an automated and reliable CT-based classification framework is required to support early identification of benign, malignant, and normal lung conditions.MethodsThis study proposes a sequential hybrid deep learning framework that integrates convolutional and recurrent neural network components for multiclass lung cancer classification. A dataset of 1,600 CT-scan images collected from multiple hospital data repositories across Bengaluru was used and divided in an approximate 70:30 ratio for training and validation. The preprocessing pipeline includes contrast enhancement using Contrast Limited Adaptive Histogram Equalization (CLAHE), morphological operations for lung segmentation, and nodule-focused masking to isolate diagnostically relevant lung regions. Data augmentation and transfer learning were applied to improve model generalization and reduce overfitting. DenseNet201 was used for feature extraction, while a bidirectional gated recurrent unit (BiGRU) module was incorporated for sequential representation learning. Hyperparameter optimization and early stopping were used to improve training stability and classification performance.ResultsThe proposed DenseNet201-BiGRU sequential hybrid architecture achieved an overall accuracy of 95.8%. The class-wise accuracies were 97.33% for benign cases, 93.33% for malignant cases, and 96.67% for normal cases. Precision, recall, and F1-score values further demonstrated that the model maintained reliable classification performance across diverse and imbalanced CT image classes.DiscussionThe results indicate that sequential hybridization of DenseNet201-based feature extraction with BiGRU-based representation learning provides an efficient, precise, and robust framework for CT-based lung cancer classification. The proposed method improves classification reliability by combining enhanced preprocessing, focused lung-region extraction, transfer learning, and recurrent modeling. However, further validation using larger, multi-center datasets and additional clinical testing is required before real-world deployment and broader diagnostic application.]]></description>
      </item>
      </channel>
    </rss>