<?xml version="1.0" encoding="utf-8"?>
    <rss version="2.0">
      <channel xmlns:content="http://purl.org/rss/1.0/modules/content/">
        <title>Frontiers in Big Data | New and Recent Articles</title>
        <link>https://www.frontiersin.org/journals/big-data</link>
        <description>RSS Feed for Frontiers in Big Data | New and Recent Articles</description>
        <language>en-us</language>
        <generator>Frontiers Feed Generator,version:1</generator>
        <pubDate>2026-10-08T05:54:32.627+00:00</pubDate>
        <ttl>60</ttl>
        <item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1866430</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1866430</link>
        <title><![CDATA[DermArtifactDB: a multi-label dataset for artifact annotation in public dermoscopic image repositories]]></title>
        <pubdate>2026-10-08T00:00:00Z</pubdate>
        <category>Data Report</category>
        <author>Vanesa Gómez-Martínez</author><author>David Chushig-Muzo</author><author>Isabel Polimón-Olabarrieta</author><author>Cristina Soguero-Ruiz</author>
        <description></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1911540</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1911540</link>
        <title><![CDATA[Reliability-aware ensemble recommendation for open-network commerce: a dual-metric framework for jointly evaluating ranking accuracy and serving correctness]]></title>
        <pubdate>2026-10-08T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Rish Praveen S</author><author>Vani Rajasekar</author>
        <description><![CDATA[Recommendation quality in open-network commerce is conventionally measured through offline ranking metrics that implicitly assume reliable upstream retrieval, stable item identities, and consistent catalog availability. In open networks such as India's Open Network for Digital Commerce (ONDC), these assumptions are routinely violated: catalogs are contributed by heterogeneous, independently operated providers, transactions follow asynchronous Beckn-style protocol lifecycles, and identifiers drift across service boundaries. This paper presents LocalMarket, a reliability-aware recommendation platform (Next.js frontend, FastAPI backend, ensemble ranking layer) that implements ONDC-pattern transaction flows in a controlled, repository-reproducible environment. We propose and validate a dual-metric evaluation framework that decouples model effectiveness—measured through Mean Average Precision (MAP@10), Recall@10 and Normalized Discounted Cumulative Gain (NDCG@10)—from serving correctness, measured through a novel Route Resolution Success Rate (RRSR) and Wishlist Consistency Rate (WCR). Using a controlled within-system ablation across Model-only, Reliability-only, and Combined configurations on fixed traffic traces, we utilize a convex-combination ensemble containing neural collaborative filtering (NCF) (utilizing only GMF embeddings for retrieval latency), graph-based, session-aware, and content-based components primarily as a vehicle for testing this framework. While this ensemble configuration improves NDCG@10 by 70.7% over a Popularity-based baseline, our primary finding validates that four targeted reliability interventions raise RRSR from 60.0 to 99.8% and WCR from 0.0 to 100.0%. Critically, the combined configuration preserves both gains simultaneously, indicating that model-level and serving-level improvements are additive rather than substitutive. The results support a systems-level thesis: in open-network commerce recommenders, ranking accuracy and serving correctness are independent, jointly necessary axes of quality, and reporting only one materially overstates the value delivered to end users. We further contribute a reusable integration failure-mode taxonomy, a constrained ensemble-weight selection protocol, and a reproducible experiment-replay procedure intended to generalize to other open-network commerce deployments.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1948398</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1948398</link>
        <title><![CDATA[Big data analytics and sentiment analysis for health misinformation detection: a systematic literature review of influencing factors, implementation challenges and a data-driven adoption framework]]></title>
        <pubdate>2026-10-05T00:00:00Z</pubdate>
        <category>Systematic Review</category>
        <author>Khurram Shahzad</author><author>Asfa Muhammed Din Javeed</author><author>Shakil Ahmad</author><author>Mujahid Latif</author><author>Kashaf Saleem</author>
        <description><![CDATA[The study aimed to identify the factors influencing big data analytics (BDA) and sentiment analysis (SA) for health misinformation detection and associated challenges for the adoption of BDA and SA across health ecosystems. The systematic literature review (SLR) methodology was applied for addressing the study's objectives. Thirty most relevant seminal studies (peer-reviewed articles and conference proceedings) were retrieved from 11 digital platforms to conduct the SLR. Results of the study showed that advanced deep learning architectures, feature engineering, sentiment intelligence, real-time analytics, and human-centered verification were key factors that influenced big data analytics and sentiment analysis based health misinformation detection. It was also identified that annotation limitations, scalability constraints, domain adaptation challenges, ethical concerns, and sustainability challenges negatively impacted the adoption of big data and sentiment driven health misinformation detection systems. The study has developed a data-driven framework to effectively adopt and sustain big data analytics and sentiment analysis to detect health misinformation.Systematic review registration:https://osf.io/wjrm6.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1995917</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1995917</link>
        <title><![CDATA[Editorial: Responsible and robust evaluation for real-world recommendation and search systems]]></title>
        <pubdate>2026-10-05T00:00:00Z</pubdate>
        <category>Editorial</category>
        <author>Alejandro Bellogín</author><author>Pablo Sánchez</author><author>Laura Sebastiá</author>
        <description></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1938137</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1938137</link>
        <title><![CDATA[Controlled knowledge updating in memory-enabled AI agents: beyond model editing and recall]]></title>
        <pubdate>2026-09-28T00:00:00Z</pubdate>
        <category>Perspective</category>
        <author>Gabriel Chavira Juárez</author><author>Eder Jahir Gonzalez Bravo</author><author>Guadalupe Esmeralda Rivera García</author><author>Javier A. Arcos Espinosa</author><author>Salvador W. Nava-Díaz</author>
        <description><![CDATA[Large language model agents that persist across sessions, tools, users, and changing environments do more than answer isolated prompts; they accumulate state. When new evidence arrives, the central question is where that change should live: transient context, external memory, tool or workflow definitions, activation states, or model parameters. Recent benchmarks compare update mechanisms and assess memory over time, while emerging architectures coordinate multiple memory types. In this Perspective, we examine how a decision layer should select among update substrates and evaluate whether its choices remain appropriate. A controlled update is not merely a successful edit or a recalled fact; it is a decision to alter the least invasive substrate sufficient for the claim's scope, expected persistence, and evidentiary strength, including when the correct action is not to write persistent state. We propose a substrate-aware view of controlled knowledge updating organized by three principles: substrate proportionality, temporal defeasibility, and auditable continuity. This framing treats knowledge updating as a longitudinal control problem and motivates evaluation criteria that include update selection, temporal consistency, interference, reversibility, efficiency, and robustness. The resulting agenda connects memory, knowledge editing, agent evolution, and benchmarking under a single practical question: how should an agent decide what to change so that future behavior improves without accumulating uncontrolled state?]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1931695</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1931695</link>
        <title><![CDATA[Anticipating user desires: predicting software product feature demand through consumer behavioral analytics]]></title>
        <pubdate>2026-09-24T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Abdullah A. Aldaeej</author>
        <description><![CDATA[In this study, we introduce the Dynamic Behavioral Feature Predictor (DBFP), a machine learning model designed to predict feature demands for software products based on consumer behavioral analytics. The model leverages advanced techniques such as Deep Convolutional Neural Networks (1D CNN) and Long Short-Term Memory (LSTM) networks to identify complex user behavior patterns. We demonstrate that DBFP achieves a high accuracy of 99%, outperforming established models such as LSTM and 1D CNN in extensive comparative experiments. Although the model shows promising performance, we note that its accuracy may be sensitive to data quality, and further research is needed to evaluate its robustness in real-world applications. DBFP excels at identifying diverse user behavior types, offering tailored recommendations for software features that align with individual user preferences. This paper highlights the potential of behavioral analytics in personalizing software development, enhancing the user experience, and improving the efficiency of feature prioritization. By bridging the fields of machine learning, big data (BD), and software personalization, DBFP lays the foundation for future advancements in user-centric software engineering. This study not only contributes to the field of predictive analytics but also opens new avenues for applying behavioral insights to software design and development.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1921037</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1921037</link>
        <title><![CDATA[Machine-assisted adaptive feedback consensus method for large-scale group decision making based on reinforcement learning]]></title>
        <pubdate>2026-09-24T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Hongyu Yu</author><author>Xuanhua Xu</author><author>Weiwei Zhang</author>
        <description><![CDATA[IntroductionTraditional consensus-reaching processes in large-scale group decision making (LSGDM) rely heavily on empirically determined feedback parameters and have difficulty balancing consensus efficiency with the preservation of expert opinions. To address this problem, this paper proposes a reinforcement learning-based machine-assisted adaptive feedback consensus method.MethodsFirst, a hesitation degree is introduced to improve the score function of hesitant fuzzy linguistic term sets (HFLTSs), thereby reducing information loss during linguistic information transformation. Second, a comprehensive relationship matrix integrating preference similarity and social trust relationships is constructed, and K-Means clustering is employed to divide large-scale expert groups. On this basis, the consensus-reaching process is modeled as a Markov decision process (MDP), and the LSGDM consensus-reaching process is modeled as a dynamic sequential decision problem. A deep deterministic policy gradient (DDPG) agent is utilized to learn feedback adjustment strategies, enabling feedback parameters to be dynamically adjusted according to the evolution state of group opinions. Meanwhile, network text data and the TF-IDF method are combined to determine attribute weights and improve decision objectivity. Finally, an emergency decision-making case of the “Beijing-Tianjin-Hebei rainstorm” is conducted to verify the proposed method.ResultsThe results show that the proposed method can achieve the preset consensus threshold within fewer discussion rounds while effectively balancing the group consensus level and expert opinion retention.Discussion/ConclusionThese findings verify the effectiveness and feasibility of the proposed method, indicating that it can provide effective decision support for adaptive feedback consensus reaching in LSGDM emergency decision-making.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1942458</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1942458</link>
        <title><![CDATA[Evaluating the reach and equity of public spaces in Mexico City using mobility data: the case of the Utopías in Iztapalapa]]></title>
        <pubdate>2026-09-23T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Ollin D. Langle-Chimal</author><author>Natalia Cadavid-Aguilar</author><author>Alberto Meouchi-Vélez</author><author>Roberto Ponce-López</author>
        <description><![CDATA[Cities invest in social infrastructure to improve well-being, but the benefits depend on who can use these spaces and who actually does. Access is therefore not only a matter of whether people live close enough to reach a facility, but also of who travels to it and from how far. We examine this realized access using high-resolution human mobility data from Location-Based Services (LBS) for the Utopías, a network of thirteen community centers offering free recreation, culture, and care in Iztapalapa, an underserved borough of 1.8 million residents in eastern Mexico City. We combine six months of mobility traces with public transit data (GTFS), a georeferenced crime registry, and a catalog of facility activities to identify the factors associated with each site's visitors. Three factors capture distinct dimensions of reach. Transit connectivity tracks the number of visitors a site draws. Facilities with more distinctive activities, such as a planetarium or a dinosaur park, show a greater increase in visitor travel distance at weekends. By contrast, a higher share of violent crime in the surrounding area is associated with more localized and lower-income visitor catchments. Together, these findings show that the reach of social infrastructure varies not only with proximity, but also with transport, the distinctiveness of what facilities offer, and local safety.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1923373</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1923373</link>
        <title><![CDATA[Humans as infrastructure: is AI-initiated steganography in social media possible?]]></title>
        <pubdate>2026-09-18T00:00:00Z</pubdate>
        <category>Opinion</category>
        <author>Daniele Ortu</author>
        <description></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1776172</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1776172</link>
        <title><![CDATA[From crisis to new routine: shifts in urban shopping mobility and socioeconomic inequality during and after the COVID-19 shock]]></title>
        <pubdate>2026-09-15T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Yilun Xu</author><author>Mohsen Bahrami</author><author>Alex Pentland</author>
        <description><![CDATA[IntroductionThis study examines the impact of the COVID-19 pandemic on revealed shopping-location patterns and their persistence during the early recovery period. It focuses on urban shopping mobility in New York City and investigates whether observed store-selection patterns returned to, or remained different from, their pre-pandemic baseline by 2021 across different socioeconomic communities.MethodsWe use large-scale mobility and place datasets together with census information to analyze visits to general merchandise and department stores in New York City. A modified Huff gravity model and a metaheuristic calibration method are used to quantify temporal shifts in store-selection patterns. We also use unsupervised learning techniques and statistical inference models to examine heterogeneity across socioeconomic communities and evaluate changes before, during, and after the COVID-19 shock.ResultsThe proposed model captures the dynamics of shopping-location decisions and temporal visit patterns under changing pandemic-era constraints. The results suggest that New Yorkers' revealed store-selection patterns changed substantially during the COVID-19 shock, with increased emphasis on store area, chain loyalty, and nearby points of interest, and reduced sensitivity to customer-store distance. These patterns did not fully revert to their 2019 baseline by 2021, although 2021 is interpreted as an early recovery and reopening period rather than a fully post-pandemic equilibrium.DiscussionThe findings highlight the societal dimension of urban shopping mobility disruptions during crises and show that mobility-based store-selection patterns vary across socioeconomic communities. The results can help urban planners, managers, and marketers better understand heterogeneous shifts in shopping mobility and adapt strategies under crisis and recovery conditions.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1763520</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1763520</link>
        <title><![CDATA[An empirical test of antecedents and perceived performance of big data analytics in telecommunication companies]]></title>
        <pubdate>2026-09-14T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Ibrahim Magboul</author><author>Diaeldin Osman</author><author>Fadi Herzallah</author><author>Alnour Nadir Osman</author><author>Dexter Gittens</author>
        <description><![CDATA[The emergence of Data concepts, sources, types, and analysis mechanisms has changed dramatically in recent years. Nowadays, big data analytics (BDA) dominates the attention of scholars and practitioners. Since this emergence, different business sectors have started deploying BDA in various endeavors to help improve decision-making and gain a competitive advantage. Recently, some organizations reported that BDA has not yielded the expected business value. In addition, researchers have focused so far on BDA usage in developed countries. Antecedents of BDA adoption have been widely researched; nevertheless, this study is among the first to collate selected antecedents and perceived performance of BDA adoption in one integrated and validated model in the telecommunication sector. Using a purposive sampling technique, thestudy deployed structural equation modeling (SEM) to analyze 201 observations. The findings reveal that six of the eight antecedents of BDA significantly impact BDA, while two showed no significant effect. Furthermore, this research finds that BDA has a positive impact on perceived performance, thus contributing to the technology adoption literature by offering valuable insights for researchers and practitioners in a developing country like Sudan.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1914133</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1914133</link>
        <title><![CDATA[Voluntary disclosure, banking stability, and AI-augmented forensic accounting: an exploratory econometric and machine-learning study of Palestinian banks]]></title>
        <pubdate>2026-09-14T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Bahaa Subhi Razia</author><author>Najwan Ibrahim Jadallah</author><author>Qasim Zureigat</author><author>Reem Khamis</author><author>Bahaa Subhi Awwad</author>
        <description><![CDATA[IntroductionInformation asymmetry between bank managers and external stakeholders is a common relationship between financial-statement fraud and banking instability.MethodsThis study combines an AI-augmented forensic accounting framework with voluntary disclosure, banking stability indicators, and exploratory machine learning (ML) approaches. The study integrates fixed-effects regression with Logistic Regression, Random Forest, XGBoost, Isolation Forest, and SHAP-based explainability using panel data from the whole population of seven banks listed on the Palestine Exchange (2019–2025).ResultsAccording to the econometric results, there is a conditional rather than a uniform relationship between voluntary disclosure and financial stability, with variation by bank size, age, and leverage. Additionally, the exploratory machine-learning analyses indicate that nonlinear approaches could help find unusual bank-year records and instability-risk patterns that are not fully captured by traditional linear models. SHAP analysis enhanced the interpretability of model classifications, and ensemble approaches outperformed Logistic Regression in cross-validation within this small sample. The machine-learning results are considered as exploratory proof-of-concept evidence rather than externally confirmed predictive outcomes due to the small sample size and lack of independently verified fraud labels.DiscussionOverall, the study shows how AI-augmented forensic accounting can enhance supervisory prioritization, instability-risk screening, and the expert assessment of anomalous observations in institutionally unstable banking contexts, thereby complementing traditional econometric analysis.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1927033</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1927033</link>
        <title><![CDATA[Algorithmic advances in smart TV content recommendation: a structured evidence-mapping review]]></title>
        <pubdate>2026-09-14T00:00:00Z</pubdate>
        <category>Review</category>
        <author>Zhe Chen</author><author>Jing He</author><author>Yuanjia Gong</author><author>Junge Liang</author><author>Jilong Li</author>
        <description><![CDATA[Smart TV recommenders operate under constraints that are less prominent on personal devices: viewing is often passive, one account may represent several viewers, direct feedback is scarce, and programmes are long and semantically rich. We mapped research published from 2015 to 2025 using a structured evidence-review protocol. The submitted bibliography comprised 122 DOI-bearing references, including two foundational sources and 120 records used in the evidence map. Crossref verified 116 records; six stable arXiv DOI records were retained. Searches across six sources yielded 2,526 database-level results. A combined deduplication and title/abstract screening stage removed 2,414 records, leaving 112 unique records for full-text assessment; all 112 full texts were retrieved, 87 records were excluded, and 25 met the eligibility criteria. These 25 studies were combined with the 95-study initial corpus to form the 120-study evidence map. Separately, we conducted an independent dual-reviewer eligibility audit of 126 records, comprising all 120 included studies and six representative boundary exclusions. Agreement was 97.6% (123/126; Cohen's κ = 0.788). Six records were excluded by both reviewers, and three discordant judgments were resolved by joint full-record review. For analysis, each study was assigned one primary technical category and one of six mutually exclusive primary functional goals; secondary technical labels captured hybrid methods. Evidence reporting was profiled separately for data transparency, reproducibility, evaluation design, external validation, and deployment validation. The final literature corpus encompasses intelligent recommendation systems based on deep learning, sequence analysis, graph theory, shared accounts, and multimodal approaches.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1917723</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1917723</link>
        <title><![CDATA[Data-driven adaptive hybrid models for exchange rate return forecasting]]></title>
        <pubdate>2026-09-11T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Olumide Sunday Adesina</author><author>Lawrence Ogechukwu Obokoh</author>
        <description><![CDATA[BackgroundExchange-rate return forecasting is challenging because financial time series may exhibit linearity, nonlinearity, regime-switching behavior, and volatility. To address these complexities, two adaptive hybrid forecasting frameworks were developed: ATW-HyF A, which dynamically combines ARIMA, SETAR, and ANN forecasts using inverse-variance weighting, and ATW-HyF B, which extends the framework by incorporating a GARCH(1,1) volatility layer to model the conditional variance of forecast errors.MethodsThe robustness of the proposed frameworks was assessed using simulated data and three foreign exchange return series: USD/NGN, EUR/USD, and GBP/USD. Forecast performance was evaluated within a rolling-origin forecasting framework and compared with competing forecasting models.ResultsThe simulation results showed that the ATW-HyF models achieved the lowest forecast errors among the competing models. However, their performance on the empirical exchange-rate series was mixed, with other models achieving lower forecast errors for some currency pairs. In particular, the ARIMA-ANN hybrid performed best for USD/NGN and EUR/USD, whereas ATW-HyF B performed best for GBP/USD.DiscussionThe findings indicate that forecasting performance depends on the characteristics of the underlying exchange-rate series rather than on model complexity alone. Combining models with complementary predictive strengths can improve forecasting performance, particularly when the data exhibit linear, nonlinear, and volatility-related patterns. The incorporation of adaptive weighting and volatility modeling therefore provides a flexible approach to exchange-rate return forecasting under evolving market conditions.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1954626</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1954626</link>
        <title><![CDATA[Editorial: Smart forecasting: deep learning and explainable AI for real-world time series prediction]]></title>
        <pubdate>2026-09-10T00:00:00Z</pubdate>
        <category>Editorial</category>
        <author>Leonard Barolli</author><author>Antonino Ferraro</author><author>Antonio Galli</author><author>Francesco Moscato</author>
        <description></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1885965</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1885965</link>
        <title><![CDATA[IOTTRUST: graph-based anomaly detection for IoT intrusion using network flow topology and community structure analysis on UNSW-NB15]]></title>
        <pubdate>2026-09-10T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Nachaat Mohamed</author><author>Hamed Taherdoost</author>
        <description><![CDATA[IntroductionConventional machine learning approaches to IoT intrusion detection treat each network flow record as an independent observation, discarding the relational structure that connects flows across source IPs, destination IPs, and subnet communities. This article presents IOTTRUST, a graph-augmented intrusion detection framework for IoT networks that constructs a directed network flow graph from the UNSW-NB15 dataset and enriches per-flow machine learning with 24 graph-derived topology features computed per source and destination IP.MethodsAll graph topology statistics are computed exclusively from the training partition and propagated to the test partition without access to test-set labels, eliminating the temporal leakage present in an earlier version of this study. The experimental design isolates the contribution of graph topology through a three-track comparison conducted on an identical 40,000-flow stratified sample with a fixed 70/30 split (random seed 42): a Flow-Only baseline using the standard 43 UNSW-NB15 features, a Graph-Only model trained exclusively on 24 graph topology features, and the full IOTTRUST Hybrid model combining 22 numeric flow features with 24 graph features (46 total).ResultsUnder this leakage-free, size-matched evaluation, Flow-Only achieves 98.98% accuracy and AUC = 0.9995 with FPR = 0.81%; Graph-Only alone reaches 97.89% accuracy and AUC = 0.9974 with FPR = 1.79%; and IOTTRUST Hybrid achieves 99.07% accuracy, AUC = 0.9996, Precision = 97.09%, Recall = 98.30%, and FPR = 0.74%, a modest but statistically significant improvement over Flow-Only (McNemar p = 0.295 on the held-out test set, paired t-test on five-fold cross-validated F1 p = 0.026) and a substantial improvement over Graph-Only alone (McNemar p < 0.001).DiscussionCommunity structure analysis on the training graph confirms that the attacker subnet (175.45.176.x, four IPs) accounts for over 90% of attack flows, exhibiting distinctive graph signatures that graph features capture directly. Feature importance analysis shows that 12 of the top 15 most important features in the leakage-free Hybrid model remain graph topology features, accounting for 60.4% of the top-15 importance mass, confirming that network relational structure continues to carry predictive information even under this stricter evaluation protocol.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1900836</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1900836</link>
        <title><![CDATA[Fine-tuning multilingual sentence transformers for low-resource skill-concept matching across Kazakh, Russian, and English]]></title>
        <pubdate>2026-09-10T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Sandugash Serikbayeva</author><author>Madina Sambetbayeva</author><author>Valiya Ramazanova</author><author>Aigerim Yerimbetova</author><author>Zhanar Lamasheva</author><author>Zhanna Sadirmekova</author><author>Ardak Batyrkhanov</author><author>Yersaiyn Mailybayev</author>
        <description><![CDATA[IntroductionMatching short, specialized skill expressions across English, Russian, and Kazakh is challenging because general-domain multilingual encoders underperform on terse, code-mixed, domain-specific phrases, particularly in the low-resource Kazakh setting.MethodsWe fine-tuned a multilingual Sentence Transformer using a staged Multiple Negatives Ranking objective on a trilingual paraphrase corpus, including Russian augmentation pairs. We evaluated skill matching and semantic similarity across languages and assessed the resulting embeddings through downstream skill-taxonomy clustering.ResultsFine-tuning preserved English performance while improving Russian and Kazakh similarity quality. Russian cosine Pearson correlation increased from 0.8125 to 0.8221, while Kazakh cosine Pearson increased from 0.5989 to 0.6050 and Kazakh dot-product Pearson from 0.4487 to 0.4912. Agglomerative clustering improved mean silhouette from 0.27126 to 0.2851 and reduced erroneous clusters from 19.74% to 13.97%.DiscussionThe results provide evidence consistent with cross-lingual transfer as an important mechanism of improvement for Kazakh. They also motivate language-specific threshold calibration and demonstrate that intrinsic similarity improvements translate into a cleaner downstream skill taxonomy.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1918383</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1918383</link>
        <title><![CDATA[A comparative study of DistilBERT and Perceiver for deceptive review detection in hospitality review data]]></title>
        <pubdate>2026-09-09T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Ioana Ximena Remeş</author><author>Felicia Mirabela Costea</author><author>Cornelia Aurora Gyorödi</author>
        <description><![CDATA[Deceptive reviews on hospitality platforms can undermine consumer trust and affect the reliability of online reputation systems. This study presents a comparative empirical evaluation of two deep learning pipelines for deceptive review detection: a DistilBERT-based pipeline fine-tuned end-to-end on review text and complemented with rating and sentiment features, and a Perceiver-based pipeline operating on fixed sentence embeddings generated by a frozen encoder. The experiments were conducted on a combined English-language corpus of 3,322 hotel reviews assembled from the Myleott benchmark corpus and a Kaggle dataset of authentic Ritz-Carlton New York reviews, comprising approximately 2,520 genuine and 800 deceptive reviews. The corpus was evaluated using stratified 80/20 train–test splits, with 2,657 reviews used for training and 665 for testing in each split, and the deep learning models were trained for 10 epochs. To assess the robustness of the findings, the evaluation was repeated across three independent random seeds and compared with majority-class and TF-IDF plus logistic regression baselines. DistilBERT achieved the most stable performance, reaching an average accuracy of 0.9348 ± 0.0014, an F1-score of 0.9569 ± 0.0013, and a Matthews correlation coefficient of 0.8240 ± 0.0012. In contrast, the original Perceiver pipeline showed a strong tendency to collapse toward the majority class, which limited its ability to identify deceptive reviews despite apparently high recall. Although class-weighted training improved Perceiver's behavior, its results remained less reliable and more variable than those obtained with DistilBERT. The balanced TF-IDF plus logistic regression baseline performed close to DistilBERT, indicating that well-configured traditional natural language processing methods remain competitive on datasets of this size. The ablation analysis further indicates that the rating and sentiment features contribute primarily to the stability of the model across different train–test splits, rather than providing a substantial independent gain in discriminative performance. The results indicate that the fine-tuned DistilBERT pipeline provides the most robust option in the evaluated setting, while also emphasizing the importance of reporting corpus composition, test-set size, class imbalance, and imbalance-aware metrics when developing deceptive review detection systems for hospitality review data.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1938279</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1938279</link>
        <title><![CDATA[AI-driven cybersecurity for industrial internet of things: architectures, challenges, datasets, and future research directions]]></title>
        <pubdate>2026-09-07T00:00:00Z</pubdate>
        <category>Review</category>
        <author>Siddhartha Singhal</author><author>Kakelli Anil Kumar</author>
        <description><![CDATA[While the Industrial Internet of Things (IIoT) has a wide range of applications in the modern era, including smart manufacturing, healthcare, transportation, energy, and critical infrastructure, the multitude of devices and distributed communication, alongside the convergence of cyber and physical systems, makes these environments vulnerable to more advanced cyber-attacks. Traditional signature or pattern-based security solutions continue to be ineffective against new, sneaky, and zero-day attack strategies, fueling the interest in AI-powered cybersecurity. This review aims to analyze the latest developments systematically in intelligent threat detection and defense in IIoT environments. The review critically analyzes the cybersecurity research published over the past few years (2020–2026) on cyber threats across the various layers of the IIoT architecture, publicly available cybersecurity datasets, evaluation practices, and AI-based intrusion detection methods, such as machine learning, deep learning, hybrid architectures, transformers, graph neural networks, federated learning, reinforcement learning, and explainable AI. High benchmark performance alone is not sufficient to claim cybersecurity effectiveness, as the synthesis shows persistent limitations in cross-domain generalization, computational overhead, explainability, adversarial robustness, edge deployment, and operational validation. Emerging research priorities included in this review are lightweight edge intelligence, continual and adaptive learning, explainable federated intelligence, digital-twin-enabled security, foundation-model-driven cyber intelligence, autonomous cyber defense, and trustworthy AI. This review provides a pathway toward resilient, adaptive, and operationally deployable cybersecurity solutions for next-generation IIoT and highlights the PRISMA-based identification process.]]></description>
      </item><item>
        <guid isPermaLink="true">https://www.frontiersin.org/articles/10.3389/fdata.2026.1876811</guid>
        <link>https://www.frontiersin.org/articles/10.3389/fdata.2026.1876811</link>
        <title><![CDATA[Reconstruction-aware urban crime analytics from incomplete judicial records using heterogeneous graph temporal imputation]]></title>
        <pubdate>2026-09-07T00:00:00Z</pubdate>
        <category>Original Research</category>
        <author>Wen Xu</author><author>Yong Dai</author><author>Qi Zhang</author><author>Keyu Chen</author><author>Sanxia Zeng</author>
        <description><![CDATA[IntroductionIncomplete judicial records constrain reliable crime analytics because missing attributes, heterogeneous case descriptions, mixed variable types, and imbalanced charge distributions reduce their analytical value. This study developed a reconstruction-aware framework for drug-crime charge classification from incomplete judicial records.MethodsDrug-related judicial documents from Chengdu, China, covering 2014-2021 were screened, cleaned, and converted into structured person-level observations. The final dataset comprised 12,620 valid documents and 15,184 observations with 16 structured features. A Heterogeneous Graph Convolution Temporal Autoencoder (HetGConv-TAE) was developed to jointly model heterogeneous case-attribute relations and temporal changes in case composition. Performance was evaluated under 10%, 20%, and 30% controlled missingness, followed by XGBoost charge classification and SHAP interpretation.ResultsHetGConv-TAE achieved the best mean reconstruction performance across conventional, static graph, temporal, and relational graph baselines. XGBoost trained on HetGConv-TAE reconstructed data achieved weighted F1 scores close to 80% and ROC-AUC values above 90%. SHAP analysis identified drug weight, place category, and administrative district as the most influential predictors across charge categories.DiscussionCombining label-excluded heterogeneous graph reconstruction, temporal encoding, downstream classification, and interpretable analysis improves the analytical usefulness of incomplete judicial archives. The framework is intended to support data-quality assessment and aggregate public-safety research rather than automated legal decision-making.]]></description>
      </item>
      </channel>
    </rss>