<?xml version="1.0" encoding="utf-8"?>
    <rss version="2.0">
      <channel xmlns:content="http://purl.org/rss/1.0/modules/content/">
        <title>Frontiers in Big Data | New and Recent Articles</title>
        <link>https://www.frontiersin.org/journals/big-data</link>
        <description>RSS Feed for Frontiers in Big Data | New and Recent Articles</description>
        <language>en-us</language>
        <generator>Frontiers Feed Generator,version:1</generator>
        <pubDate>2026-09-15T23:09:42.985+00:00</pubDate>
        <ttl>60</ttl>
        <item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1776172</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1776172</link>
        <title><![CDATA[From crisis to new routine: shifts in urban shopping mobility and socioeconomic inequality during and after the COVID-19 shock]]></title>
        <pubdate>2026-09-15T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Yilun Xu</author><author>Mohsen Bahrami</author><author>Alex Pentland</author>
        <description><![CDATA[IntroductionThis study examines the impact of the COVID-19 pandemic on revealed shopping-location patterns and their persistence during the early recovery period. It focuses on urban shopping mobility in New York City and investigates whether observed store-selection patterns returned to, or remained different from, their pre-pandemic baseline by 2021 across different socioeconomic communities.MethodsWe use large-scale mobility and place datasets together with census information to analyze visits to general merchandise and department stores in New York City. A modified Huff gravity model and a metaheuristic calibration method are used to quantify temporal shifts in store-selection patterns. We also use unsupervised learning techniques and statistical inference models to examine heterogeneity across socioeconomic communities and evaluate changes before, during, and after the COVID-19 shock.ResultsThe proposed model captures the dynamics of shopping-location decisions and temporal visit patterns under changing pandemic-era constraints. The results suggest that New Yorkers' revealed store-selection patterns changed substantially during the COVID-19 shock, with increased emphasis on store area, chain loyalty, and nearby points of interest, and reduced sensitivity to customer-store distance. These patterns did not fully revert to their 2019 baseline by 2021, although 2021 is interpreted as an early recovery and reopening period rather than a fully post-pandemic equilibrium.DiscussionThe findings highlight the societal dimension of urban shopping mobility disruptions during crises and show that mobility-based store-selection patterns vary across socioeconomic communities. The results can help urban planners, managers, and marketers better understand heterogeneous shifts in shopping mobility and adapt strategies under crisis and recovery conditions.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1927033</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1927033</link>
        <title><![CDATA[Algorithmic advances in smart TV content recommendation: a structured evidence-mapping review]]></title>
        <pubdate>2026-09-14T00:00:00Z</pubdate>
        <category>Review</category>
        <author>Zhe Chen</author><author>Jing He</author><author>Yuanjia Gong</author><author>Junge Liang</author><author>Jilong Li</author>
        <description><![CDATA[Smart TV recommenders operate under constraints that are less prominent on personal devices: viewing is often passive, one account may represent several viewers, direct feedback is scarce, and programmes are long and semantically rich. We mapped research published from 2015 to 2025 using a structured evidence-review protocol. The submitted bibliography comprised 122 DOI-bearing references, including two foundational sources and 120 records used in the evidence map. Crossref verified 116 records; six stable arXiv DOI records were retained. Searches across six sources yielded 2,526 database-level results. A combined deduplication and title/abstract screening stage removed 2,414 records, leaving 112 unique records for full-text assessment; all 112 full texts were retrieved, 87 records were excluded, and 25 met the eligibility criteria. These 25 studies were combined with the 95-study initial corpus to form the 120-study evidence map. Separately, we conducted an independent dual-reviewer eligibility audit of 126 records, comprising all 120 included studies and six representative boundary exclusions. Agreement was 97.6% (123/126; Cohen's κ = 0.788). Six records were excluded by both reviewers, and three discordant judgments were resolved by joint full-record review. For analysis, each study was assigned one primary technical category and one of six mutually exclusive primary functional goals; secondary technical labels captured hybrid methods. Evidence reporting was profiled separately for data transparency, reproducibility, evaluation design, external validation, and deployment validation. The final literature corpus encompasses intelligent recommendation systems based on deep learning, sequence analysis, graph theory, shared accounts, and multimodal approaches.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1763520</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1763520</link>
        <title><![CDATA[An empirical test of antecedents and perceived performance of big data analytics in telecommunication companies]]></title>
        <pubdate>2026-09-14T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Ibrahim Magboul</author><author>Diaeldin Osman</author><author>Fadi Herzallah</author><author>Alnour Nadir Osman</author><author>Dexter Gittens</author>
        <description><![CDATA[The emergence of Data concepts, sources, types, and analysis mechanisms has changed dramatically in recent years. Nowadays, big data analytics (BDA) dominates the attention of scholars and practitioners. Since this emergence, different business sectors have started deploying BDA in various endeavors to help improve decision-making and gain a competitive advantage. Recently, some organizations reported that BDA has not yielded the expected business value. In addition, researchers have focused so far on BDA usage in developed countries. Antecedents of BDA adoption have been widely researched; nevertheless, this study is among the first to collate selected antecedents and perceived performance of BDA adoption in one integrated and validated model in the telecommunication sector. Using a purposive sampling technique, thestudy deployed structural equation modeling (SEM) to analyze 201 observations. The findings reveal that six of the eight antecedents of BDA significantly impact BDA, while two showed no significant effect. Furthermore, this research finds that BDA has a positive impact on perceived performance, thus contributing to the technology adoption literature by offering valuable insights for researchers and practitioners in a developing country like Sudan.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1914133</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1914133</link>
        <title><![CDATA[Voluntary disclosure, banking stability, and AI-augmented forensic accounting: an exploratory econometric and machine-learning study of Palestinian banks]]></title>
        <pubdate>2026-09-14T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Bahaa Subhi Razia</author><author>Najwan Ibrahim Jadallah</author><author>Qasim Zureigat</author><author>Reem Khamis</author><author>Bahaa Subhi Awwad</author>
        <description><![CDATA[IntroductionInformation asymmetry between bank managers and external stakeholders is a common relationship between financial-statement fraud and banking instability.MethodsThis study combines an AI-augmented forensic accounting framework with voluntary disclosure, banking stability indicators, and exploratory machine learning (ML) approaches. The study integrates fixed-effects regression with Logistic Regression, Random Forest, XGBoost, Isolation Forest, and SHAP-based explainability using panel data from the whole population of seven banks listed on the Palestine Exchange (2019–2025).ResultsAccording to the econometric results, there is a conditional rather than a uniform relationship between voluntary disclosure and financial stability, with variation by bank size, age, and leverage. Additionally, the exploratory machine-learning analyses indicate that nonlinear approaches could help find unusual bank-year records and instability-risk patterns that are not fully captured by traditional linear models. SHAP analysis enhanced the interpretability of model classifications, and ensemble approaches outperformed Logistic Regression in cross-validation within this small sample. The machine-learning results are considered as exploratory proof-of-concept evidence rather than externally confirmed predictive outcomes due to the small sample size and lack of independently verified fraud labels.DiscussionOverall, the study shows how AI-augmented forensic accounting can enhance supervisory prioritization, instability-risk screening, and the expert assessment of anomalous observations in institutionally unstable banking contexts, thereby complementing traditional econometric analysis.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1917723</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1917723</link>
        <title><![CDATA[Data-driven adaptive hybrid models for exchange rate return forecasting]]></title>
        <pubdate>2026-09-11T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Olumide Sunday Adesina</author><author>Lawrence Ogechukwu Obokoh</author>
        <description><![CDATA[BackgroundExchange-rate return forecasting is challenging because financial time series may exhibit linearity, nonlinearity, regime-switching behavior, and volatility. To address these complexities, two adaptive hybrid forecasting frameworks were developed: ATW-HyF A, which dynamically combines ARIMA, SETAR, and ANN forecasts using inverse-variance weighting, and ATW-HyF B, which extends the framework by incorporating a GARCH(1,1) volatility layer to model the conditional variance of forecast errors.MethodsThe robustness of the proposed frameworks was assessed using simulated data and three foreign exchange return series: USD/NGN, EUR/USD, and GBP/USD. Forecast performance was evaluated within a rolling-origin forecasting framework and compared with competing forecasting models.ResultsThe simulation results showed that the ATW-HyF models achieved the lowest forecast errors among the competing models. However, their performance on the empirical exchange-rate series was mixed, with other models achieving lower forecast errors for some currency pairs. In particular, the ARIMA-ANN hybrid performed best for USD/NGN and EUR/USD, whereas ATW-HyF B performed best for GBP/USD.DiscussionThe findings indicate that forecasting performance depends on the characteristics of the underlying exchange-rate series rather than on model complexity alone. Combining models with complementary predictive strengths can improve forecasting performance, particularly when the data exhibit linear, nonlinear, and volatility-related patterns. The incorporation of adaptive weighting and volatility modeling therefore provides a flexible approach to exchange-rate return forecasting under evolving market conditions.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1900836</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1900836</link>
        <title><![CDATA[Fine-tuning multilingual sentence transformers for low-resource skill-concept matching across Kazakh, Russian, and English]]></title>
        <pubdate>2026-09-10T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Sandugash Serikbayeva</author><author>Madina Sambetbayeva</author><author>Valiya Ramazanova</author><author>Aigerim Yerimbetova</author><author>Zhanar Lamasheva</author><author>Zhanna Sadirmekova</author><author>Ardak Batyrkhanov</author><author>Yersaiyn Mailybayev</author>
        <description><![CDATA[IntroductionMatching short, specialized skill expressions across English, Russian, and Kazakh is challenging because general-domain multilingual encoders underperform on terse, code-mixed, domain-specific phrases, particularly in the low-resource Kazakh setting.MethodsWe fine-tuned a multilingual Sentence Transformer using a staged Multiple Negatives Ranking objective on a trilingual paraphrase corpus, including Russian augmentation pairs. We evaluated skill matching and semantic similarity across languages and assessed the resulting embeddings through downstream skill-taxonomy clustering.ResultsFine-tuning preserved English performance while improving Russian and Kazakh similarity quality. Russian cosine Pearson correlation increased from 0.8125 to 0.8221, while Kazakh cosine Pearson increased from 0.5989 to 0.6050 and Kazakh dot-product Pearson from 0.4487 to 0.4912. Agglomerative clustering improved mean silhouette from 0.27126 to 0.2851 and reduced erroneous clusters from 19.74% to 13.97%.DiscussionThe results provide evidence consistent with cross-lingual transfer as an important mechanism of improvement for Kazakh. They also motivate language-specific threshold calibration and demonstrate that intrinsic similarity improvements translate into a cleaner downstream skill taxonomy.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1954626</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1954626</link>
        <title><![CDATA[Editorial: Smart forecasting: deep learning and explainable AI for real-world time series prediction]]></title>
        <pubdate>2026-09-10T00:00:00Z</pubdate>
        <category>Editorial</category>
        <author>Leonard Barolli</author><author>Antonino Ferraro</author><author>Antonio Galli</author><author>Francesco Moscato</author>
        <description></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1885965</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1885965</link>
        <title><![CDATA[IOTTRUST: graph-based anomaly detection for IoT intrusion using network flow topology and community structure analysis on UNSW-NB15]]></title>
        <pubdate>2026-09-10T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Nachaat Mohamed</author><author>Hamed Taherdoost</author>
        <description><![CDATA[IntroductionConventional machine learning approaches to IoT intrusion detection treat each network flow record as an independent observation, discarding the relational structure that connects flows across source IPs, destination IPs, and subnet communities. This article presents IOTTRUST, a graph-augmented intrusion detection framework for IoT networks that constructs a directed network flow graph from the UNSW-NB15 dataset and enriches per-flow machine learning with 24 graph-derived topology features computed per source and destination IP.MethodsAll graph topology statistics are computed exclusively from the training partition and propagated to the test partition without access to test-set labels, eliminating the temporal leakage present in an earlier version of this study. The experimental design isolates the contribution of graph topology through a three-track comparison conducted on an identical 40,000-flow stratified sample with a fixed 70/30 split (random seed 42): a Flow-Only baseline using the standard 43 UNSW-NB15 features, a Graph-Only model trained exclusively on 24 graph topology features, and the full IOTTRUST Hybrid model combining 22 numeric flow features with 24 graph features (46 total).ResultsUnder this leakage-free, size-matched evaluation, Flow-Only achieves 98.98% accuracy and AUC = 0.9995 with FPR = 0.81%; Graph-Only alone reaches 97.89% accuracy and AUC = 0.9974 with FPR = 1.79%; and IOTTRUST Hybrid achieves 99.07% accuracy, AUC = 0.9996, Precision = 97.09%, Recall = 98.30%, and FPR = 0.74%, a modest but statistically significant improvement over Flow-Only (McNemar p = 0.295 on the held-out test set, paired t-test on five-fold cross-validated F1 p = 0.026) and a substantial improvement over Graph-Only alone (McNemar p < 0.001).DiscussionCommunity structure analysis on the training graph confirms that the attacker subnet (175.45.176.x, four IPs) accounts for over 90% of attack flows, exhibiting distinctive graph signatures that graph features capture directly. Feature importance analysis shows that 12 of the top 15 most important features in the leakage-free Hybrid model remain graph topology features, accounting for 60.4% of the top-15 importance mass, confirming that network relational structure continues to carry predictive information even under this stricter evaluation protocol.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1918383</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1918383</link>
        <title><![CDATA[A comparative study of DistilBERT and Perceiver for deceptive review detection in hospitality review data]]></title>
        <pubdate>2026-09-09T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Ioana Ximena Remeş</author><author>Felicia Mirabela Costea</author><author>Cornelia Aurora Gyorödi</author>
        <description><![CDATA[Deceptive reviews on hospitality platforms can undermine consumer trust and affect the reliability of online reputation systems. This study presents a comparative empirical evaluation of two deep learning pipelines for deceptive review detection: a DistilBERT-based pipeline fine-tuned end-to-end on review text and complemented with rating and sentiment features, and a Perceiver-based pipeline operating on fixed sentence embeddings generated by a frozen encoder. The experiments were conducted on a combined English-language corpus of 3,322 hotel reviews assembled from the Myleott benchmark corpus and a Kaggle dataset of authentic Ritz-Carlton New York reviews, comprising approximately 2,520 genuine and 800 deceptive reviews. The corpus was evaluated using stratified 80/20 train–test splits, with 2,657 reviews used for training and 665 for testing in each split, and the deep learning models were trained for 10 epochs. To assess the robustness of the findings, the evaluation was repeated across three independent random seeds and compared with majority-class and TF-IDF plus logistic regression baselines. DistilBERT achieved the most stable performance, reaching an average accuracy of 0.9348 ± 0.0014, an F1-score of 0.9569 ± 0.0013, and a Matthews correlation coefficient of 0.8240 ± 0.0012. In contrast, the original Perceiver pipeline showed a strong tendency to collapse toward the majority class, which limited its ability to identify deceptive reviews despite apparently high recall. Although class-weighted training improved Perceiver's behavior, its results remained less reliable and more variable than those obtained with DistilBERT. The balanced TF-IDF plus logistic regression baseline performed close to DistilBERT, indicating that well-configured traditional natural language processing methods remain competitive on datasets of this size. The ablation analysis further indicates that the rating and sentiment features contribute primarily to the stability of the model across different train–test splits, rather than providing a substantial independent gain in discriminative performance. The results indicate that the fine-tuned DistilBERT pipeline provides the most robust option in the evaluated setting, while also emphasizing the importance of reporting corpus composition, test-set size, class imbalance, and imbalance-aware metrics when developing deceptive review detection systems for hospitality review data.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1938279</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1938279</link>
        <title><![CDATA[AI-driven cybersecurity for industrial internet of things: architectures, challenges, datasets, and future research directions]]></title>
        <pubdate>2026-09-07T00:00:00Z</pubdate>
        <category>Review</category>
        <author>Siddhartha Singhal</author><author>Kakelli Anil Kumar</author>
        <description><![CDATA[While the Industrial Internet of Things (IIoT) has a wide range of applications in the modern era, including smart manufacturing, healthcare, transportation, energy, and critical infrastructure, the multitude of devices and distributed communication, alongside the convergence of cyber and physical systems, makes these environments vulnerable to more advanced cyber-attacks. Traditional signature or pattern-based security solutions continue to be ineffective against new, sneaky, and zero-day attack strategies, fueling the interest in AI-powered cybersecurity. This review aims to analyze the latest developments systematically in intelligent threat detection and defense in IIoT environments. The review critically analyzes the cybersecurity research published over the past few years (2020–2026) on cyber threats across the various layers of the IIoT architecture, publicly available cybersecurity datasets, evaluation practices, and AI-based intrusion detection methods, such as machine learning, deep learning, hybrid architectures, transformers, graph neural networks, federated learning, reinforcement learning, and explainable AI. High benchmark performance alone is not sufficient to claim cybersecurity effectiveness, as the synthesis shows persistent limitations in cross-domain generalization, computational overhead, explainability, adversarial robustness, edge deployment, and operational validation. Emerging research priorities included in this review are lightweight edge intelligence, continual and adaptive learning, explainable federated intelligence, digital-twin-enabled security, foundation-model-driven cyber intelligence, autonomous cyber defense, and trustworthy AI. This review provides a pathway toward resilient, adaptive, and operationally deployable cybersecurity solutions for next-generation IIoT and highlights the PRISMA-based identification process.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1876811</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1876811</link>
        <title><![CDATA[Reconstruction-aware urban crime analytics from incomplete judicial records using heterogeneous graph temporal imputation]]></title>
        <pubdate>2026-09-07T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Wen Xu</author><author>Yong Dai</author><author>Qi Zhang</author><author>Keyu Chen</author><author>Sanxia Zeng</author>
        <description><![CDATA[IntroductionIncomplete judicial records constrain reliable crime analytics because missing attributes, heterogeneous case descriptions, mixed variable types, and imbalanced charge distributions reduce their analytical value. This study developed a reconstruction-aware framework for drug-crime charge classification from incomplete judicial records.MethodsDrug-related judicial documents from Chengdu, China, covering 2014-2021 were screened, cleaned, and converted into structured person-level observations. The final dataset comprised 12,620 valid documents and 15,184 observations with 16 structured features. A Heterogeneous Graph Convolution Temporal Autoencoder (HetGConv-TAE) was developed to jointly model heterogeneous case-attribute relations and temporal changes in case composition. Performance was evaluated under 10%, 20%, and 30% controlled missingness, followed by XGBoost charge classification and SHAP interpretation.ResultsHetGConv-TAE achieved the best mean reconstruction performance across conventional, static graph, temporal, and relational graph baselines. XGBoost trained on HetGConv-TAE reconstructed data achieved weighted F1 scores close to 80% and ROC-AUC values above 90%. SHAP analysis identified drug weight, place category, and administrative district as the most influential predictors across charge categories.DiscussionCombining label-excluded heterogeneous graph reconstruction, temporal encoding, downstream classification, and interpretable analysis improves the analytical usefulness of incomplete judicial archives. The framework is intended to support data-quality assessment and aggregate public-safety research rather than automated legal decision-making.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1925931</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1925931</link>
        <title><![CDATA[HK-DeepIV: heat-kernel geometric diagnostics and early-warning signals for curvature-induced interference in financial correlation networks]]></title>
        <pubdate>2026-09-03T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Ntebogang Dinah Moroke</author>
        <description><![CDATA[Financial correlation networks under infrastructure shocks undergo curvature-driven interference that standard estimators cannot detect. This study introduces Heat-Kernel Deep Instrumental Variables (HK-DeepIV), a geometric diagnostic and early-warning framework modeling shock propagation as heat diffusion on the Riemannian manifold of Johannesburg Stock Exchange (JSE) asset-return correlations. Three formal results underpin the architecture: a non-parametric identification result for the projection of the structural function onto the leading heat-kernel eigenfunction (k = 1; a scalar instrument can identify at most one linear combination of the eigenfunctions, and no claim is made beyond that projection); double robustness via Neyman orthogonality; and Corollary 3.2, showing that unit-level distinguishability collapses exponentially in diffusion time whenever Ollivier-Ricci curvature is positive; a purely geometric statement providing the basis for the fragility diagnostics developed here. Applied to 60 JSE tickers (2,832 trading days, 2015–2025) with Eskom load-shedding as the treatment (2022–2025), the trained model produces an ATE of +383 bp, reported as an overfitting artifact. Five independent baselines converge on smaller, mostly negative estimates. The most credibly conditioned conditional-association estimate is double machine learning [−20.75 bp, heteroskedasticity-and-autocorrelation-consistent (HAC)-corrected 95% CI [−51.62, 10.12] bp, p = 0.188], which is not significant at conventional levels; consistent with the instrument exogeneity caveat (5-day lagged return balance test, p = 0.007) and the descriptive framing of all estimates. The instrument [48-h-ahead Eskom stage forecast, Corr(Zt, Dt)≈0.71, first-stage F = 47.3] is distinct from the treatment (realized Stage ≥2 binary), but exogeneity is not confirmed; all estimates are conditional associations. The Fiedler eigenvalue (mean 0.3827, minimum 0.1819) and mean Ollivier-Ricci curvature (0.5517, persistently positive) provide computable real-time fragility indicators. Across the four most severe load-shedding quarters, Fiedler and Ricci diagnostics offer comparable early-warning signals (mean lead-time difference −0.5 trading days); neither is systematically superior, but their combination is more informative than either alone. The framework contributes toward real-time spectral monitoring of infrastructure-driven systemic risk, supporting Uited Nations Sustainable Development Goal (UN SDG 9) in emerging markets.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1901121</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1901121</link>
        <title><![CDATA[HyperIDS: a hypergraph learning and quantum-inspired transformer ensemble framework for IoT intrusion detection]]></title>
        <pubdate>2026-09-01T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Samayank Goel</author><author>Logeswari Govindaraj</author><author>Tamilarasi Kathirvel Murugan</author>
        <description><![CDATA[The unprecedented increase in the number of Internet of Things (IoT) devices has widened the attack surface of modern-day networks, making them vulnerable to various cyberattacks. Conventional intrusion detection mechanisms face challenges in identifying the subtle correlations between network traffic characteristics and providing consistent detection rates in varying heterogeneous environments. With this context, this research aims at introducing HyperIDS, an innovative Intrusion Detection System (IDS) which combines Hypergraph Learning, Quantum-Inspired Feature Selection and Optimization, and Transformer-based Ensemble Classifier for intelligent IoT cyberattack detection. First, Dual-Fitness Enhanced Gaussian Quantum Particle Swarm Optimization (DFE-GQPSO) approach is utilized to select the most relevant traffic characteristics while reducing feature space and computational cost. These selected traffic features are then converted to a hypergraph form, allowing higher order relations between different entities of the network to be captured. Following this step, Hypergraph Neural Networks (HGNN) is deployed to generate structural and relational representations from the hypergraph structure of the dataset. Long-term dependencies and attack patterns are subsequently extracted using a transformer encoder. The final classification process involves combining CatBoost and XGBoost using stacking ensemble method. In addition, a SHAP-based explainability module is incorporated to ensure transparency and trustworthiness of the developed system. In order to evaluate the proposed framework, HyperIDS is experimentally tested against two commonly used cyber security datasets, namely, TON_IoT and Bot-IoT. Experimental findings have shown superior effectiveness of HyperIDS in detecting cyber-attacks with 98.92, 98.81, 98.76, and 98.78% of accuracy, precision, recall, and F1 score, respectively on the TON_IoT dataset. Similarly, accuracy, precision, recall, and F1 scores of HyperIDS reach 99.14, 99.05, 99.01, and 99.03% on the Bot-IoT dataset. Comparison with conventional machine learning (ML), deep learning (DL), and hybrid intrusion detection techniques have proven HyperIDS's superiority in detecting cyberattacks on IoT infrastructure.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1924721</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1924721</link>
        <title><![CDATA[GICPIdb: an archival repository of multimodal data focusing on pathological images for gastrointestinal cancers]]></title>
        <pubdate>2026-08-31T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Huang Chen</author><author>Ling Tong</author><author>Jinyang Liu</author><author>Kai Liu</author><author>Xintao Li</author><author>Shufang Shi</author><author>Shuxue Xi</author><author>Geng Tian</author><author>Meijun Zhang</author><author>Dingrong Zhong</author><author>Shijun Li</author><author>Jialiang Yang</author>
        <description><![CDATA[IntroductionDeep learning (DL) shows great potential for predicting biomarkers from routine histopathological slides of gastrointestinal (GI) cancers. Yet most existing models are validated on limited patient cohorts, while pathological image annotation and molecular marker standardization demand substantial professional expertise. To address these gaps, we constructed the Gastrointestinal Cancer Pathological Image Archive (GICPIdb, gicpidb.shubuzuo.top), a dedicated database and web platform covering seven major GI cancer types.MethodsHigh-quality hematoxylin and eosin (H&E)-stained whole-slide images were collected from multiple sources and uniformly processed. Image annotations were performed by board-certified pathologists following standardized protocols. GICPIdb offers five interactive web modules for data uploading, quality control, feature extraction, online annotation and AI-based prediction. Its intuitive interface supports data browsing, retrieval, visualization and downloading.ResultsThe database houses 2,863 pathologist-annotated, uniformly processed, high-quality H&E stained images collected from 2,655 patients. Of these, 1,699 patients were sourced from The Cancer Genome Atlas (TCGA), 182 from the Clinical Proteomic Tumor Analysis Consortium (CPTAC), and 424 from China-Japan Friendship Hospital and 350 from Chifeng Municipal Hospital in Inner Mongolia, China. It also integrates data on over 50 key molecular markers (e.g., MSI, TMB) and prognostic labels related to survival, recurrence and metastasis.DiscussionGICPIdb aims to promote the development of DL-driven AI tools for cancer research and clinical translation. The multi-institutional data collection and standardized annotation pipeline are expected to enhance the generalizability and reproducibility of AI-based prediction models across diverse patient populations.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1888885</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1888885</link>
        <title><![CDATA[Study of deep learning cues for cross linguistic part of speech tagging in English– Malayalam code-mixed data]]></title>
        <pubdate>2026-08-28T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Parvathy Padmakumar</author><author>Shreya S. Nair</author><author>A. K. Prajisha</author><author>S. Thara</author>
        <description><![CDATA[IntroductionPart Of Speech (POS) tagging is a fundamental task in Natural Language Processing (NLP) that assigns grammatical labels to words in a sentence. Code mixed text, which entails switching between two or more languages within a single conversation or a sentence, presents challenges for POS tagging. This investigation entailed a comprehensive study of deep learning approaches for cross linguistic POS tagging, focused on English Malayalam code mixed data prevalent on social media platforms. The study was carried out on linguistically complex and varied English Malayalam code mixed text from social media platforms with informal spellings, language switching, transliteration, slang, and unclear grammatical boundaries, reflecting the characteristics of informal online communication. We provide the first structured analysis of micro macro divergence and rare class behavior in English Malayalam code mixed POS tagging.MethodsWe evaluated 14 state of the art model configurations that span traditional sequence labeling approaches and multilingual transformer architectures. Models were compared using standard performance metrics prevalent in the domain of data science, supplemented by normalized confusion matrices, error prone tag identification and micro-macro F1 gap analysis.ResultsOur results showed that CRF (No Lang) emerged as the most balanced model overall on macro F1 (all classes) of 0.8170, while (BiLSTM + CRF) achieved the highest macro F1 (seen classes) of 0.8831, precision of 0.9167, and recall of 0.875, though this reflects strong performance concentrated on frequent tag classes rather than balanced coverage across the full tag set. Notably, the pretrained multilingual transformers (mBERT, MuRIL), despite prior exposure to Malayalam during pretraining, were outperformed on several key metrics by CRF and BiLSTM models trained directly on the code mixed dataset.DiscussionThis finding was contrary to our expectation that existing multilingual knowledge would translate into a clear advantage on this task and merits further investigation.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1863692</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1863692</link>
        <title><![CDATA[BRI DataLab: an AI-assisted research platform for Chinese infrastructure finance]]></title>
        <pubdate>2026-08-27T00:00:00Z</pubdate>
        <category>Technology and Code</category>
        <author>Sanoop Sajan Koshy</author>
        <description><![CDATA[BRI DataLab is an open-source web platform that integrates a harmonized infrastructure financing dataset, a curated policy and academic document corpus, and a retrieval-augmented generation (RAG) interface into a single queryable research environment for Chinese overseas development finance. Analysts working on the Belt and Road Initiative (BRI) must move between datasets, government documents, and academic literature without an integrated interface for querying across all three at once. This paper documents BRI DataLab, a research platform that combines a harmonized project-level dataset derived from AidData's China's Global Loans and Grants Dataset v1.0 with a curated corpus of 42 policy documents, institutional reports, and academic sources, accessed through an AI-assisted retrieval and synthesis layer built on RAG. Three methodological contributions are described: the decisions involved in converting more than 33,000 tranche-level financing records into 4,861 infrastructure projects; the four-category selection framework governing the document collection; and the system design of the AI interface, including the epistemic constraints built into its system prompt. The central argument is that AI-assisted research interfaces are methodologically defensible when they are transparent about what they can and cannot establish.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1897881</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1897881</link>
        <title><![CDATA[A novel three-parameter Gompertz-Lomax Distribution with a physics-informed survival network: statistical theory, hazard shape characterization, and real-world applications]]></title>
        <pubdate>2026-08-27T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Furqan Ahmed P.</author><author>Sujatha V.</author>
        <description><![CDATA[This article introduces the Gompertz-Lomax Distribution (GLD), a novel three-parameter continuous lifetime model whose composite hazard rate additively combines an increasing Gompertz component with a decreasing Lomax component. The resulting closed-form survival function nests the Gompertz and Lomax distributions as exact special cases and accommodates increasing, decreasing, and bathtub-shaped hazard profiles within a single parametric framework. Complete closed-form expressions are derived for the raw and central moments, incomplete moments, quantile function, mean residual life, stress-strength reliability, stochastic ordering, Rényi entropy, and Bonferroni-Lorenz inequality curves. Four estimation procedures, namely maximum likelihood, method of moments, least squares, and weighted least squares, are developed and validated through a Monte Carlo simulation study comprising N = 10, 000 replications, with finite-sample performance compared across all four estimators. As the primary machine-learning contribution, a Physics-Informed Survival Network (PISN) is introduced: a deep neural network that maps subject-level covariates to individual GLD shape and scale parameters while encoding the theoretically derived hazard shape condition as a smooth, differentiable penalty in the training loss. Empirical evaluation on three benchmark datasets, namely, leukemia patient survival, COVID-19 hospital survival, and electronic component failure times, demonstrates that the GLD achieves among the lowest Kolmogorov-Smirnov statistics and the most favorable Akaike Information Criterion (AIC), corrected AIC (AICC), and Bayesian Information Criterion (BIC) on all three datasets, with ΔAIC exceeding five units over the nearest competitor on the leukemia dataset and ΔAIC = 30.6 on the electronic component dataset, which exhibits a bathtub-shaped hazard rate.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1889044</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1889044</link>
        <title><![CDATA[A cognitively inspired feature-level fusion framework for interpretable retail sales forecasting using integration of extreme gradient boost machine, artificial neural network and attention mechanism model]]></title>
        <pubdate>2026-08-20T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Munienge Mbodila</author><author>Omobayo Ayokunle Esan</author>
        <description><![CDATA[IntroductionAccurate retail sales forecasting in modern data-intensive environments requires models that not only achieve high predictive accuracy but also scale efficiently as data volume and feature complexity increase. This study proposes a novel cognitively inspired hybrid framework, XGB–ANN–Attn, for interpretable and scalable retail analytics.MethodsThe model introduces feature-level fusion by integrating XGBoost-derived leaf embeddings with deep neural representations, enabling joint modeling of low-order statistical dependencies and high-order nonlinear feature interactions. A lightweight attention mechanism dynamically assigns instance-specific feature importance, enhancing both predictive performance and interpretability while maintaining linear computational complexity with respect to feature dimensionality. Unlike conventional ensemble approaches that operate at the decision level, the proposed framework integrates representations at the level of representations, reducing redundancy and improving computational efficiency. The model is designed for scalability, combining the log-linear complexity of gradient boosting with the linear scaling properties of neural networks and attention mechanisms.ResultsExperimental evaluation on the BigMart and Walmart datasets demonstrates superior performance compared to state-of-the-art models, achieving RMSE = 0.1584 and R2 = 0.9946 on BigMart, and RMSE = 0.8652 and R2 = 0.9568 on Walmart. Furthermore, the framework supports parallelization and distributed deployment, making it suitable for large-scale retail systems.DiscussionThe alignment between attention weights and SHAP explanations provides transparent and actionable insights. The results confirm that the proposed approach offers a scalable, interpretable, and high-performance solution for AI-driven decision support systems in big-data retail environments.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1812391</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1812391</link>
        <title><![CDATA[Interviewer effects on underreporting of traditional contraceptive use: evidence from National Family Health Survey-5]]></title>
        <pubdate>2026-08-20T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Rupalee Singh Chauhan</author><author>Laxmi Kant Dwivedi</author><author>Priyanka Dixit</author><author>Shiva S. Halli</author>
        <description><![CDATA[BackgroundSocioeconomic alignment between interviewers and respondents may play a critical role in shaping the accuracy of reported reproductive histories. This study investigates how interviewer characteristics contribute to the reporting of adoption of traditional contraceptive methods using reproductive event calendar data from the fifth round of the National Family Health Survey (2019–21) in India.Data and methodsUsing contraceptive calendar data, a cross-classified multilevel model was employed to assess the interviewer's influence on the reporting of traditional contraceptive method use. The analysis included women aged 15–49 years who reported contraceptive use in the calendar month(s), and the interviewers interviewed them.ResultsMultilevel cross-classified models show that interviewer characteristics significantly influence reporting of traditional contraceptive use following periods of non-use, childbirth, or pregnancy termination. Women interviewed by older interviewers (≥30 years; adjusted odds ratio (AOR) = 1.40, p < 0.01) and those who shared the same religion (AOR = 1.21, p ≤ 0.01) and marital status (AOR = 1.42, p < 0.01) as the interviewer were more likely to report traditional contraceptive method use. While women of the same age group (AOR = 0.89, p < 0.05), same-state residence (AOR = 0.82, p < 0.01), and same caste (AOR = 0.94, p < 0.01) as the interviewer were less likely to report traditional contraceptive method use. Older interviewers (AOR = 0.86, p < 0.01), those with secondary education (AOR = 0.54, p < 0.01), those belonging to the Muslim religion (AOR = 0.43, p < 0.01), and those with a lack of prior experience in conducting surveys (AOR = 0.90, p < 0.10) were associated with lower reporting. Moderate (AOR = 1.35, p < 0.01) and high (AOR = 1.22, p < 0.05) levels of prior experience with the same survey schedule in conducting the survey significantly improved reporting. Distribution of variance across different levels indicates substantial variation at the interviewer level (28.50%), followed by the district (18.50%) and PSU (13.46%) levels, suggesting that interviewer effects play a significant role in shaping the outcome.Discussion and conclusionInterviewer characteristics such as education, age, and prior survey experience influence the reporting of traditional contraceptive use. Similarities or differences between interviewers and respondents in characteristics such as age, marital status, religion, caste, and state of residence also affect the accuracy and completeness of reporting traditional contraceptive use in contraceptive calendar data.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1927115</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1927115</link>
        <title><![CDATA[A novel two-parameter discrete exponentiated exponential distribution with its neutrosophic extension: mathematical characterization and real-world count data modeling]]></title>
        <pubdate>2026-08-18T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Furqan Ahmed P.</author><author>Sujatha V.</author>
        <description><![CDATA[This paper proposes and studies the Discrete Exponentiated Exponential (DEE) distribution, a new two-parameter discrete lifetime model obtained by applying the survival discretization technique to the exponentiated exponential distribution. The DEE distribution possesses a key structural advantage: the single shape parameter α simultaneously controls both the hazard rate shape (increasing for 0 < α < 1, decreasing for α > 1) and the dispersion regime (over-dispersion for small α, under-dispersion for large α), a dual flexibility that most competing discrete models cannot replicate without structural modification. These properties are rigorously established in closed form via the Glaser technique and numerical investigation of the dispersion index. Closed-form expressions are derived for the probability mass function, cumulative distribution function, survival and hazard rate functions, moment generating and characteristic functions, raw and central moments up to the fourth order, and the mean residual life function. Parameter estimation is carried out via maximum likelihood, and a comprehensive Monte Carlo simulation study assesses the finite-sample performance of the estimators and benchmarks it against four competing discrete models under identical settings. The practical advantage of the DEE distribution is demonstrated through three real-world count datasets from oncology, education, and nephrology; the DEE model exhibits competitive and frequently superior performance relative to seven benchmark discrete distributions. Finally, a neutrosophic generalization (DNEE) extends the framework to count data characterized by vagueness, incompleteness, or indeterminacy, with accompanying estimation and illustrative numerical analysis.]]></description>
      </item>
      </channel>
    </rss>