Abstract
Background:
Pathway enrichment analyses are widely used to interpret transcriptomic datasets; however, their outputs typically consist of lists of statistically enriched pathways that require qualitative interpretation and are difficult to compare across biological contexts. Methods of semantic classification that transform enrichment results into quantitative, mechanistically interpretable measures of system-level dysregulation remain underexplored.
Methods:
Here, we introduced TENSE (quanTifiEd immuNe-aging dySregulation index), a framework that summarizes pathway enrichment outputs into a quantitative estimate of immune-aging–associated dysregulation. Utilizing a Large Language Model classifier via a KNIME workflow, significantly enriched pathways are semantically classified into five mechanistic categories representing key processes implicated in immune aging, the DIRES scheme: DNA damage (D), DNA repair (R), epigenetic drift (E), inflammaging (I), and nucleic acid sensing (S). These pathway-derived signals are then aggregated into a normalized dysregulation score reflecting the magnitude (TENSE) and distribution (DIRES) of aging-associated processes across biological contexts.
Results:
Application of TENSE to transcriptional modules derived from neurodegenerative, radiation-response, and immune activation datasets revealed distinct dysregulation profiles. Alzheimer’s disease–associated modules were primarily characterized by inflammaging signatures, particularly within microglial transcriptional programs, whereas radiation response datasets exhibited dominant DNA damage-related signals. Sepsis-associated gene signatures showed strong inflammatory contributions, producing the highest TENSE values observed. Robustness analysis demonstrated high reproducibility of pathway classification across repeated runs and close agreement between large language model–derived annotations and human consensus scores.
Conclusion:
TENSE provides a reproducible and interpretable method for transforming pathway enrichment outputs into quantitative estimates of system-level immune-aging dysregulation. By bridging pathway enrichment analysis and mechanistic interpretation, the framework enables comparative analysis of aging-related biological processes across diverse datasets.
1 Introduction
Temporal accounting of biological processes is an essential aspect of system coordination for living organisms, from the cellular level (). Biological time for an organism can be defined as a set of iteratively organized biological processes that maintain cellular and organismal homeostasis throughout its lifespan. Biological time is inescapably linked with physical time (). As organisms progress along physical time, the fidelity of biological processes declines in the human body, defining the process of aging (Tenchov et al., 2024). The loss of biological process fidelity reflects the decay of all cellular systems, from nuclear processes to cellular metabolism and transcellular communication ().
One of the more prominent drivers of aging is the progressive accumulation of DNA damage, reflecting a decline in the DNA damage response efficiency (). However, declining repair efficacy leads to the accumulation of additional damage in a feed-forward manner (Schumacher et al., 2021). As a central process in aging, unmitigated DNA damage initiates another feedforward process: inflammation mediated by nucleic acid sensors. This induction, along with potential causes of dysregulated mitigation processes, drives age-related inflammation (Zhao et al., 2023). The loss of genomic stability due to the failure of repair mechanisms thus marks a salient point of no return, where time-dependent disorder prevails over its mitigation (; Vavougios et al., 2022; ).
In addition to the idea of inflammaging (), these concepts suggest an immune arrow of time—a biological clock that indicates the gradual decline in the accuracy of essential homeostatic processes, observed, as increasing inflammatory tonicity.
The implied parallel between additive disorder in biological systems and aging can also be used to describe the critical role of inflammation in the maturation and development of stem cells (Tatullo, 2024), as well as the critical role of the inflammatory milieu for tissue-specific processes such as synaptogenesis and neurogenesis (Paul et al., 2021). In short, the immune arrow of time may also accelerate purposefully due to physiological processes, as well as secondary to pathophysiological conditions.
An informative approach to identifying systems-level dysregulation is transcriptomic profiling to extract disease-associated gene signatures (; ; ) and the corresponding enriched pathways (; Wang et al., 2026; Song et al., 2026). Gene signature methods quantify the coordinated activity of gene networks, generating weighted scores that can predict outcomes in specific disease contexts (Sparks et al., 2024). Similarly, pathway enrichment analyses identify biological processes that are statistically overrepresented within a dataset, enabling the detection of pathways potentially involved in the underlying biological condition (Vavougios et al., 2022). Similarly, current state-of-the-art applications of LLMs focus on extracting knowledge from gene sets and annotating their function (Wang et al., 2025). While these approaches provide valuable insights into molecular activity and pathway involvement, they primarily address specific research questions within defined biological contexts. Consequently, a methodological gap remains for metrics capable of integrating system-level genomic signals into a quantitative estimate of the magnitude and distribution of dysregulation across predefined biological axes.
In this study, we propose an alternative strategy for annotating pathway enrichment results to extract biologically meaningful information on dysregulation and disease-associated processes. As a proof of concept, we introduce the TENSE (quanTifiEd immuNe-aging dySregulation indEx) scoring framework, designed to quantify immune aging–associated regulatory disorder in biological systems. TENSE functions as an integrative, normalized, systems-level metric that summarizes pathway-derived dysregulation into a single interpretable index, enabling comparative assessment of immune-aging dynamics across datasets and experimental contexts.
TENSE is designed to transform pathway enrichment outputs into a structured representation of immune-aging dynamics by mapping enriched pathways to a set of mechanistically defined biological axes. These axes correspond to key processes implicated in immune aging, namely DNA damage (D), DNA repair (R), epigenetic drift (E), inflammaging (I), and nucleic acid sensing (S), which together form the DIRES scoring scheme. Recent advances in artificial intelligence and computational biology have enabled increasingly sophisticated interpretation of high-dimensional data. These applications include machine learning–based modeling, pathway analysis, and network-level integration approaches. Broadly, these developments illustrate the capacity of AI-driven methods to extract structured insights from complex datasets across diverse scientific domains (Raza and Hanif, 2026; Raza et al., 2026; ; Raza et al., 2025a; Raza et al., 2025b).
Large language models (LLMs) are increasingly integrated into bioinformatics workflows, including assisting with knowledge curation of the Reactome database (Wu et al., 2025). LLMs have been shown to retrieve semantic knowledge to predict regulatory relationships based on provided gene networks, albeit with varying performance ().
In the TENSE workflow, we capitalize LLM capabilities for semantic analysis and classification of enriched pathways into mechanistic aging axes; we then employ them to compute a systems-level dysregulation metric. The LLM core serves as a semantic classifier within this framework, rather than functioning as a biological database. It interprets pathway descriptions and categorizes them into established mechanistic groups: DNA damage, DNA repair, epigenetic drift, inflammaging, and nucleic acid sensing. Pathway nomenclature often contains biological significance that is not consistently captured by simple keyword matching or static database mapping. For instance, pathways such as “HDR through MMEJ” or “OAS antiviral response” necessitate contextual interpretation to appropriately associate them with DNA repair or nucleic acid-sensing processes. The integration of an LLM semantic processing layer, GPT-4.1-NANO, enables scalable and reproducible semantic classification of enriched pathways while retaining mechanistic clarity. It is important to note that the LLM layer GPT-4.1-NANO does not generate biological conclusions; its role is limited to pathway annotation, while the TENSE metric is calculated according to defined quantitative rules. This rigidity also serves to develop TENSE in a relatively controlled data environment while the intended core functionalities are assessed.
TENSE combines the magnitude of enrichment and the pathway contributions along each axis to deliver a normalized quantitative assessment of immune-aging–related dysregulation. With this approach, pathway enrichment outcomes are viewed not simply as lists of significant pathways, but as an integrated metric that represents both the extent and spread of biological disruption among aging-associated processes.
We explore TENSE scoring by reanalyzing relevant published studies and procuring several cases, with a primary focus on Alzheimer’s disease. We demonstrate the utility of the TENSE and its associated measures in annotating pathway enrichment data and the biological context of each studied condition.
2 Methods
2.1 Theoretical framework
To define TENSE, we consider factors that drive the accumulation of disorder in a biological system as a function of age, herein a cell or tissue, building toward incremental genomic instability (). For the development of this model, major drivers of aging, adjacent to genomic instability, were considered and modeled as dimensionless variables; Specifically, we considered the equilibrium between DNA damage and repair (; Schumacher et al., 2021; ; Vlachogiannis et al., 2023; Whittemore et al., 2019), nucleic acid sensing mechanisms (Paul et al., 2021; Stillman et al., 2024; ; ; ; Taffoni et al., 2021), inflammaging (; ) and epigenetic drift (; ; ; ).
Although the relationships among these biological processes are complex and involve multiple feedback and regulatory interactions, for the purpose of computing TENSE, their contributions to cumulative dysregulation are modeled as additive. This assumption reflects both the simplifying representation of system-level dysregulation and the statistical framework of pathway enrichment analysis, in which biological pathways are evaluated as independent enrichment events (Nguyen et al., 2019). Consequently, enriched pathways are treated as independent proxies for the underlying biological processes implicated in disease pathogenesis.
Thus, we define the Immune-Entropy Quotient as a function of aging-related processes:
Where:
: DNA damage as a function of time.
: Inflammaging as a function of time.
: DNA repair efficiency as a function of time, with R(t) > 0.
: Epigenetic drift as a function of time.
: Nucleic acid sensing as a function of time.
Each of α,β,γ,δ: weighting parameters reflecting the relative contribution of each term.
The normalized Immune-Entropy quotient would then be calculated with the following formula, when :
To calculate normalized H(t) from enrichment datasets, we replace proxy scores derived from gene set overrepresentation analyses. To reformulate the function using pathway proxies, we consider the following selection criteria:
Proxy pathways for reflect equilibrium between damage and repair, as gene set data typically do not contain DNA damage metadata. Thus, we consider enriched pathways additively rather than as a ratio of damage to response.
Proxy pathways for will include those related to viral infection or intracellular parasites; this choice is intended to capture enriched pathways that may not necessarily reflect infection, but the dysregulation of nucleic acid-sensing/pattern recognition receptors, as reflected by the overrepresented pathways.
Proxy pathways for will include those related to over-represented epigenetic biological functions.
Proxy pathways for will include those related to inflammatory signaling cascades.
A notable modeling convention is that, although the conceptual framework regards DNA damage and repair as interconnected processes (D/R), the pathway-based implementation treats them in an additive manner. This approach aligns with the methodology of pathway enrichment analysis, in which enrichment results are generated from distinct pathway annotations, offering independent signals rather than direct assessments of counteracting activities. As a result, pathways associated with damage and repair may demonstrate significant enrichment independently when analyzing a specific gene set. The TENSE (quanTifiEd immuNe-aging dySregulation index) is reformulated as follows:
Where:
represent the arithmetic means of the combined scores obtained via enrichment analysis (Xie et al., 2021) for each proxy pathway for DNA damage, DNA repair, Epigenetic Drift, Inflammaging, and Nucleic Acid Sensing, correspondingly, after weighting for multiple memberships. Specifically, to calculate , i ∈ D, I, R, E, S, while accounting for multiple memberships (i.e., a term belonging to more than one of the DIRES categories), we apply a 1/k weight per combined score for each term, where k = the number of categories each term is a member of. As an example, a pathway identified as belonging to the D, R, S categories (k = 3) would have its combined score divided by 3.
represents a normalized sum of pathway proxy counts for each of DNA damage, DNA repair, Epigenetic Drift, Inflammaging and Nucleic Acid Sensing, correspondingly. To calculate , i ∈ D, I, R, E, S the null hypothesis accepted is that the probability of an enriched pathway is a proxy for at least one of these five categories is 20% (i.e., one in five). The 20% reference probability was used as a neutral null model arising from the five predefined DIRES categories. Under this assumption, in the absence of category-specific enrichment, an enriched pathway assigned to the DIRES framework is expected to have equal prior probability of mapping to any one of the five categories. This assumption is not intended to imply equal biological prevalence of the five processes, but rather to provide a transparent normalization baseline.
If the null hypothesis stands, is initially calculated as the total number of pathway proxies divided by [0.2(Total Number of Enriched Pathways) This formulation produces results that can be interpreted as fold-enrichment, i.e., indicating over- or under-representation based on the ratio observed against expected:
indicates “as expected under the 20% baseline assumption of the model”.
indicates “over-represented”.
“under-represented”.
We again constrain by applying a 1/k for each D, I, R, E, S count, to account for multiple memberships in one of the five DIRES categories, leading to standardized values between 0 and 1 for the calculation of TENSE. For example, a pathway scoring in D, R, and S would have a count of 1.
4 If the null hypothesis is rejected, i.e., an enriched pathway does not belong to DIRES, = 0, i.e., it does not contribute to TENSE dysregulation.
To provide a standardized estimate of the Entropy Quotient, henceforth the TENSE metric, we divide by .
6 Standardized TENSE as calculated in step 4 provides values ranging from:
A minimum of 0, when no enriched pathway belongs to DIRES.
A maximum of 5, when each of , after applying the 1/k weight for the five DIRES categories.
To obtain normalized TENSE, producing a range of values between 0 and 1, we divide by the maximum TENSE = 5. Implementation of TENSE and DIRES scoring via an LLM-powered KNIME workflow.
To calculate TENSE in an automated manner, a hybrid semantic – deterministic workflow is designed based on the previously presented theoretical framework and implemented via an LLM-powered KNIME analytical workflow.1 The overall architecture uses pathway enrichment analysis via Enrichr to produce combined enrichment scores that adjusts for potential bias in favoring larger length gene sets in gene set libraries (). Thus, Enrichr combined scores provide weighted contributions to pathways by incorporating both statistical significance and enrichment magnitude. Notably, while Enrichr combined scores provide a useful measure of pathway relevance within a given analysis, they are not inherently comparable across datasets. The TENSE framework addresses this limitation by normalizing pathway contributions, enabling comparison of the relative structure of dysregulation across biological contexts. Within the TENSE framework, these scores are utilized to maintain the relative importance of pathways within enrichment analyses, followed by normalization across DIRES categories. Consequently, TENSE characterizes the distribution of dysregulation in each dataset rather than depending on direct comparison of raw enrichment scores from independent analyses.
Pathways included in each DIRES category were selected for their established mechanistic links to biological processes relevant to immune aging, such as genomic instability, DNA repair capacity, epigenetic regulation, chronic inflammatory activation, and nucleic acid sensing. Enrichment analyses utilized Enrichr,2 which offers curated pathway annotations derived from recognized biological databases including Reactome, KEGG, and Gene Ontology. These curated pathways enable enriched terms to be mapped to mechanistically interpretable biological processes, a requirement of the DIRES framework underlying the TENSE metric. Conversely, clustering approaches like MCODE (Zhou et al., 2019) in Metascape identify densely connected gene clusters based on network topology. However, while useful for exploratory network analysis, these clusters are dependent on individual datasets and may not align with predefined biological mechanisms.
Because the TENSE framework requires consistent mechanistic axes across datasets to quantify immune-aging–associated dysregulation, curated pathway annotations were preferred over topology-based clustering methods. Subsequently, enriched terms and combined scores are then analyzed by an LLM-powered KNIME workflow to provide DIRES and TENSE scoring. KNIME is an open-source workflow management platform with a graphical user interface that enables users to design and implement sophisticated data science pipelines (Song et al., 2026). In recent updates, KNIME has incorporated AI nodes to facilitate the integration of large language models, further empowering designed workflows. Figure 1 presents an overview of the study design.
Figure 1
2.2 Semantic layer: LLM-powered scoring of enriched biological processes
For the semantic layer of the workflow, natural language processing of enriched processes was performed via an API node using GPT-4.1-nano. GPT-4 has been previously utilized for processing the Reactome DB (Tiwari et al., 2023) and biomedical literature (Poretsky et al., 2025), and thus the GPT-4.1-nano3 was selected. GPT-4.1-nano is one of the default choices provided by the OpenAI LLM node in KNIME. The main advantages it provides are cost-effectiveness and speed according to its specifications.4 The use of GPT-4.1-nano is intended as a proof of concept and relies on at least one use case involving GPT-4 processing of Reactome data. The application of other LLMs, newer versions of GPT-4.1-nano (e.g., GPT-5-nano), as well as manual pathway annotation and TENSE calculation, is equally feasible.
In KNIME, GPT-4.1-nano is accessed via a unique API key, which is applied to an authenticator node and connected to an OpenAI LLM selector node. For this proof-of-concept application, GPT-4.1-nano was used with default settings, with a sampling temperature of 0.2 and a maximum response length of 200 tokens. Notably, sampling temperatures typically range from 0.0 to 1 (and above), with higher values allowing for creative responses, whereas lower temperatures, e.g., 0.0 to 0.3, correspond to answers with greater precision and are more appropriate when exact answers are required ().
To score enriched pathways by DIRES, LLM prompter nodes were connected serially to GPT-4.1-nano and subsequently prompted via Expression Nodes to score significantly enriched terms (provided by input nodes; defined by a false discovery rate <0.05). The prompts are available in the online version of the workflow and from Supplementary File 1. An example of a D node prompt is the following, delivered to an LLM prompter node via a KNIME Expression Node:
“You are a strict binary classifier.
OUTPUT CONTRACT (MANDATORY):
- Output exactly N lines, where N = number of input terms.
- Each line MUST be exactly ONE character: either 0 or 1.
- No spaces, no punctuation, no text.
- Multi-digit outputs like 10, 01, 00 are INVALID.
If any rule is violated, correct it before responding.
SELF-CHECK (DO NOT PRINT):
Verify: (1) line count == N, (2) each line is exactly one character and is 0 or 1.
TASK:
Return 1 ONLY if the pathway term directly reflects DNA damage, genotoxic stress, or genomic instability (damage presence/consequences). Else 0.
DEFINITION (DNA damage):
Processes indicating DNA lesions, genome instability, or cellular responses that signal damage presence (NOT the repair execution mechanisms).
POSITIVE ANCHORS (strong evidence to score as 1):
- DNA damage
- genomic instability
- genotoxic stress
- DNA lesion(s)
- replication stress (as damage/stress)
- DNA strand break(s) (if not explicitly ‘repair’)
- p53-mediated damage response / p53 signaling in response to damage
- ATM/ATR checkpoint signaling in response to DNA damage (signaling/checkpoint, not repair execution)
- oxidative DNA damage / ROS-induced DNA damage
OVERRIDE RULE (ALWAYS 1 if present — highest priority):
If the term contains any of these tokens (case-insensitive), output 1:
DNA damage, genotoxic, genomic instability, replication stress, DNA lesion, strand break, double-strand break (DSB) (unless explicitly ‘repair’), single-strand break (SSB) (unless explicitly ‘repair’), p53 (when clearly damage-related), ATM, ATR (when clearly damage/checkpoint related).
NEGATIVE ANCHORS (strong evidence to score as 0):
- DNA repair mechanisms
- RNA damage response mechanisms
- Nucleic acid sensing mechanisms
- Immune processes not directly triggered by DNA damage
- General cell cycle/replication without damage context
If uncertain, output 0.
INPUT TERMS (one term per line):”
The output provided at the end of this process is a table containing the DIRES scores before weighting for multi-category membership (Table 1).
Table 1
| Term | D_count | Ι_count | R_count | E_count | S_count |
|---|---|---|---|---|---|
| Immune System | 0 | 0 | 0 | 0 | 0 |
| Innate Immune System | 0 | 1 | 0 | 0 | 0 |
| Neutrophil Degranulation | 0 | 1 | 0 | 0 | 0 |
| Keratinization | 0 | 0 | 0 | 0 | 0 |
| Cytokine Signaling in the Immune System | 0 | 1 | 0 | 0 | 0 |
| Signaling by Interleukins | 0 | 1 | 0 | 0 | 0 |
| RAC1 GTPase Cycle | 0 | 0 | 0 | 0 | 0 |
An example of DIRES output following serial processing and scoring.
The terms are Reactome pathways, whereas the count columns score binary for DNA damage (D_count), Inflammaging (I_count), DNA Repair (R_count), Epigenetic drift (E_count), and Nucleic Acid Sensing (S_count).
This scoring scheme is performed independently for each D, I, R, E, S category and serially, D ➔ I ➔R ➔E ➔S, using a dedicated LLM prompter node and prompt. This design choice was intended to provide greater precision for each category while providing a dissectible architecture for troubleshooting and further development. The final DIRES scoring table is then submitted to the deterministic layer for further calculations.
2.3 Deterministic layer: calculating TENSE via DIRES
The deterministic layer utilizes multiple KNIME nodes to calculate TENSE. To improve interpretation, we also provide normalized values for TENSE and for each DIRES category. This design choice enables easier interpretation as TENSE can score from 0 to 1 following normalization (i.e., 0–100% dysregulation), rather than 0–5 as a standardized value. To normalize TENSE’s values, we divide the derivative score from steps 1–5 by 5, i.e., the maximum standardized value of TENSE.
Specifically, the following step-wise procedure is followed to calculate TENSE from the DIRES pivot table.
Accepts input from the Semantic Layer in the form of the DIRES pivot table, containing the scores of each term in the D, I, R, E, S categories.
Calculates the sum of each D, I, R, E, S via Math Formula KNIME nodes.
Calculates the arithmetic mean of the combined scores (provided by Enrichr) for each term scoring in D, I, R, E, S for each category via Math Formula KNIME nodes.
Calculates the standardized D, I, R, E, S scores by dividing the sum of proxy pathways in each category by [0.2(Total Number of Enriched Pathways)].
Calculate TENSE as a standardized score by applying 1/k weighting as previously described.
As an overview of the workflow, the following process is followed to calculate DIRES and TENSE:
LLM-based scoring of enriched pathways for membership in D, I, R, E, S.
Calculation of the standardized D, I, R, E, S scores by dividing the number of proxy pathways in each category by [0.2(Total Number of Enriched Pathways)].
Calculation of the arithmetic means of the combined scores for D, I, R, E, and S proxies.
Calculate the standardized TENSE by .
Normalize TENSE score to a range of values between 0,1.
When none of the DIRES categories have enriched pathways, all components of the TENSE score become zero, resulting in a zero denominator. In the KNIME workflow, this leads to a missing value. To ensure clarity and consistency, these cases are deliberately assigned a TENSE value of 0. This approach indicates that there is no DIRES-associated dysregulation in the analyzed pathways, instead of representing a numerical calculation from non-zero values. Therefore, a TENSE score of 0 should be understood as an absence of immune-aging–related dysregulation within the boundaries of this framework.
2.4 Quality control and accuracy
2.4.1 Eidetic layer: restricting the black box effect of the LLM-scoring process
For quality control and to reduce black-box effects, we established a parallel explanatory layer that provides insight into the LLM scoring process and decision-making. This explanatory layer uses the same prompts, albeit instructed to provide a rationale rather than a binary classifier exclusively. This layer can be connected to the main classifier to provide insight into the DIRES scoring process.
As an example, for the first term in Table 1, the Eidetic Layer provides the following exchange when checking the term “Immune System” for membership in D:
2.4.2 GPT-4.1-nano response
The term “Immune System” is primarily associated with the body’s defense mechanisms against pathogens, including bacteria, viruses, and other foreign substances. While the immune system can sometimes be involved in processes related to DNA damage—such as immune responses to DNA damage or immune cells in recognizing damaged cells—it is not inherently or exclusively a DNA damage process.
Therefore, since “Immune System” broadly pertains to immune responses rather than directly to DNA damage processes, the appropriate response is: **0**.
**Rationale:** The immune system’s primary functions are immune defense, immune regulation, and pathogen recognition, not DNA repair or damage processes.
We subsequently assess the rationale for each answer to ensure consensus between manual scoring by the author and GPT-4.1-nano. Notably, this example, “Immune System,” is a broad term that does not specify processes that map to inflammaging and, by default, is not intended to score in the I category to avoid artificial inflation.
2.5 Accuracy and precision metrics with reference to a human consensus
2.5.1 Theoretical approach
To address LLM non-determinism, i.e., the variability between repeated responses for the same prompt (Shusterman et al., 2025), we performed multiple (n = 100) sequential runs of a single dataset with at least one valid term for each category in DIRES.
We then extract TENSE values for each run and calculate mean TENSE±SD and 95% Confidence interval (CI). Subsequently, we score the dataset using DIRES as a human-consensus and provide a human-consensus TENSE value.
Subsequently, we use the latter value to calculate the following metrics:
Bias of the Mean (BM): Mean TENSE (LLM) − Human Consensus TENSE (Wilmoth, 2012; ).
2.5.2 Implementation dataset
To assess TENSE’s accuracy and precision, we utilize a transcriptomic module from Sparks et al. (2024) study. In brief, Sparks et al. (2024) performed weighted gene co-expression network analyses (WGCNA) of whole-blood transcriptomics from 228 patients of 22 monogenic immune diseases and 42 age- and sex-matched healthy participants. WGCNA identified 12 transcriptional modules that are differentially enriched for immune processes and cells. Among available TMs, we select TM1, which is enriched for Type I interferon responses. The rationale for this selection is that Type I interferon signaling and related pathways are expected to map broadly to DIRES categories (; Papatriantafyllou, 2013). This will ensure that a non-zero TENSE score can be calculated and used to identify oscillations caused by LLM stochasticity.
TENSE is calculated as described above via a looped version of the TENSE workflow, utilizing the Counting Loop Start and Loop End nodes for n = 100 iterations of the workflow; These workflow iterations furthermore correspond to 5 prompt iterations for each of DIRES categories per significantly enriched pathway (npathway = 38), for a total of 100x5x38 = 1,540 total prompt iterations utilizing GPT-4.1-nano via LLM prompter nodes. (Supplementary File 1). TENSE scores produced via this process (Mean TENSE LLM = 0.242 ± 0.015, 95% CI:0.240–0.244) demonstrated high concordance with human consensus (BM = −0.006; RBM = −2.7%) and stable performance across repeated runs (SD = 0.015; CV = 4.24%; Table 2). The TENSE values produced by the workflow across 100 runs are available as Supplementary material 1 (Figure 2).
Table 2
| Metric | Value |
|---|---|
| Reference TENSE | 0.248 |
| LLM mean TENSE | 0.242 |
| Standard deviation | 0.015 |
| 95% confidence interval (Half-Width) | 0.002 |
| Bias of the mean (BM) | −0.006 |
| Relative BM | −0.027 |
| Coefficient of variation (CV) | 0.042 |
TENSE calculations following 100 iterations of GPT-4.1-nano.
Reference TENSE is calculated by a human consensus, i.e., by annotating each pathway as DNA damage, Inflammaging, DNA repair, epigenetic drift, and then manually calculating TENSE. LLM: large language model, here GPT-4.1-nano.
Figure 2
Subsequently, we leveraged Reactome Pathways from TM10, a transcriptional module enriched for housekeeping cell signaling pathways, which, by design, are not expected to score in DIRES. The workflow loop produced a consistent TENSE value of 0 across 100 independent workflow iterations, indicating no spurious DIRES activation and demonstrating high classifier specificity.
2.5.3 Interoperability and modularity
To ensure TENSE’s application beyond the manual version of the workflow, we implemented Anthropic’s Claude Haiku 4.5 (Available as CLAUDE-HAIKU-4-5-20251001 from the KNIME LLM nodes) as an alternative LLM classifier for the semantic layer at the same sampling temperature T = 0.2, as previously described. The Haiku model was selected as relatively equivalent to GPT-4.1-nano in terms of cost efficiency for API calls, accuracy, and speed (Available from: https://platform.claude.com/docs/en/about-claude/models/overview; Accessed on March 9th, 2026).
TM9, enriched for RNA and RNA-related processes, was utilized as the comparison dataset. Genes comprising the module underwent the process described previously to produce a Reactome term list and associated Enricher-derived metrics.
Using fixed prompts for both LLM classifiers, Claude Haiku produced comparable results (Table 3 and Figure 3), indicating that the TENSE / DIRES workflow is interoperable between LLMs.
Table 3
| Reactome term | mODEL | D_rlt | I_rlt | R_rlt | E_rlt | S_rlt | TENSE |
|---|---|---|---|---|---|---|---|
| RNA metabolism | GPT-4.1-nano | 0.044 | 0.013 | 0.033 | 0.013 | 0.002 | 0.025 |
| RNA metabolism | claude-haiku-4-5-20251001 | 0.022 | 0.030 | 0.041 | 0.009 | 0.001 | 0.041 |
Head-to-head comparison between GPT4.1-nano and Claude Haiku 4.5.
The “_rlt” suffix denotes the utilization of normalized relative scores (i.e., , i ∈ D, I, R, E, S normalized bounded by 0,1) for each DIRES category.
Figure 3
Claude Haiku produced several formatting errors that skewed DIRES; specifically, instead of binary classification, response cells contained text with descriptive answers. An example follows:
Input: Pathway to be assessed: PIP3 Activates AKT Signaling:
Response: Here are the results for the terms in “PIP3 Activates AKT Signaling”: PIP3: 0, Activates: 0, AKT: 0, Signaling: 0
For the calculation of TENSE using Claude Haiku 4.5, formatting errors were treated as missing values; GPT-4.1-nano did not produce format errors. Despite fail-safes and consistent performance by GPT-4.1-nano, these format errors in Claude Haiku 4.5 response structure may plausibly reflect the result of prompt structure and its bias toward GPT-4.1-nano nativity. Further refinement with Claude was not pursued beyond the feasibility phase, in favor of GPT-4.1-nano’s performance on the given task and, by extension, API cost-efficiency.
To assess the robustness of the workflow in scoring a predetermined biological concept, we implemented an alternative/perturbed DIRES scheme where:
DNA damage and DNA repair mechanisms are aggregated in the D node.
Cellular signaling replaces DNA Repair in the R node.
Cellular metabolism replaces Epigenetic in the E node.
We then reanalyzed TM10, enriched for cellular signaling, and with a zero TENSE score. The perturbed TENSE score was 0.246 (24.6%), with cellular signaling cascades inflating the novel R node. We conclude that the workflow is adequately modular in the sense that it can serve biological concepts beyond those intended by TENSE’s design, e.g., function as a cell signaling/metabolic/stress classifier based on the design of the semantic layer.
2.6 Access and availability
The current version of the TENSE classifier workflow is available via the KNIME community hub under/gvavou. An individual API key, or manual DIRES annotation of the datasets to be analyzed, is required to run each iteration of the workflow.
3 Results
3.1 Use cases for TENSE and DIRES annotation in the literature
Following the construction of the TENSE engine, we examine the utility of both TENSE and DIRES scores by annotating published datasets in various settings as use cases. We examine the TENSE/DIRES annotations in prior studies to determine whether they adequately capture biological information and offer further insight, as well as a novel means of presenting that information. Raw data usage from referenced studies, Enrichr access to Reactome,5 KNIME workflow utilization,6 and API Calls to GPT-4.5-nano (via KNIME AI integration) were performed on March 9th, 2026.
3.1.1 Use case 1: annotating Alzheimer’s disease immune aspects on neuron and microglial gene signatures
3.1.1.1 Background and rationale
For the first use case, we will annotate cell-specific gene signatures associated with Alzheimer’s disease diagnosis and pathology. We will furthermore examine the concordance between Neuron- and microglia-specific gene signatures across two distinct studies (; ). Specifically, we utilized a Neuronal-derived consensus gene module from medial temporal gyri from Alzheimer’s disease compared to control, as reported by . Briefly, Chen and colleagues performed weighted gene co-expression network analysis (WGCNA) on 10,000 highly variable genes, identifying eight co-expression gene modules, four of which (yellow, brown, pink, and turquoise) showed to correlate with AD pathology.
DIRES and TENSE annotations on Chen et al.’s dataset were compared to annotations on corresponding neuron-enriched (Consensus Module 1; CM1) and microglia-enriched (Consensus Module 8; CM8) AD-associated gene modules reported by . In their study, reported on WCGNA-derived consensus gene modules that were correspondingly microglia- and neuron-specific, and associated with Alzheimer’s disease.
Finally, we aimed to determine whether TENSE and DIRES annotations would be conceptually concordant with data reported by by comparing the response transcriptomes between Αβ plaques vs. neurofibrillary tangles, which revealed inflammation as the most salient component.
3.1.1.2 Approach and results
We performed enrichment analysis for each pair of gene modules using Enrichr and identified significantly over-represented pathways from Reactome. We replicated the steps described previously. Subsequently, we obtain the following values for TENSE:
For the AD-related modules by Chen and colleagues, we obtain the following values: TENSEyellow = 0, TENSEpink = 0, TENSEbrown = 0.131 (13.1%), and TENSEturquoise = 0.012 (1.2%).
For the neuronal module (CM1), we obtain TENSECM1 = 0.022 (2.2%) and for the microglial module (CM8) TENSECM8 = 0.208 (20.8%) (Figures 4–7).
Figure 4
Figure 5
Figure 6
Figure 7
3.1.2 Use case 2: annotating the differential transcriptomic response of radiation-resistant and radiation-sensitive cells exposed to 4 Gy radiation at 4- and 12-h post-exposure
3.1.2.1 Background and rationale
Radiation induces both genotoxic stress and inflammation in affected cells; contrary to annotated data from AD, data from cellular irradiation inform us on the temporally resolved effects of a well-defined exposure on the transcriptome. To approach this research question, we utilized data available from . Their study reported differences in gene expression between radiosensitive (RS) and radioresistant (RR) cell lines at various time points after irradiation at a dose of 4 Gy.
3.1.2.2 Approach and results
To obtain DIRES and TENSE annotations, we replicate the steps described previously. Subsequently, we obtain the following values for TENSE:
For the radiosensitive line, TENSE = 0.044 (4.4%) at 4 h vs. TENSE = 0.364 (36.4%) at 12 h.
For the radioresistant line, TENSE = TENSE = 0.2 (20%) at both 4 h and 12 h.
We then calculate the differences between baseline and post-exposure TENSE scores.
We define the difference as ΔTENSE = TENSEpost-exposure - TENSEbaseline for both scenarios.
In the Radiosensitive line, ΔTENSE = 32%, whereas in the radioresistant line, ΔTENSE = 0.
Figures 8, 9 display Radar plots for DIRES scores for the radiosensitive and radioresistant lines, respectively. DIRES annotation suggests that the radiosensitive line exhibits an inflammatory component at 4 h, which is absent at 12 h post-exposure. Conversely, the radioresistant line’s transcriptome indicates invariant activation of repair mechanisms at both time points, and the absence of a detectable inflammatory component.
Figure 8
Figure 9
DIRES annotation outlines a sustained inflammatory and DNA damage component for radiosensitive cells, whereas radioresistant cells display an invariant DNA repair component. Differences in TENSE scores indicated a more dynamic utilization of cellular processes in radiosensitive cells and a robust yet invariant utilization for radioresistant cells, polarized toward repair.
These findings are in agreement with the findings of the original study, where differences in DNA repair capacity and mitigation of radiation-induced inflammation Yahyapour et al. (2018) between the two cell types was outlined. With our approach, both TENSE and DIRES serve to quantify these changes beyond the study of differentially expressed genes or pathways of interest. The concept of ΔTENSE, i.e., the difference between TENSE values at baseline and post-exposure, is useful for distinguishing between stable and dynamic processes.
3.1.3 Use case 3: TENSE and DIRES as descriptors of immune gene network signatures
3.1.3.1 Background and rationale
Conserved gene networks have previously been linked to impaired regulation of the host response to infection, and their activity is often captured by distinct gene signatures that link gene activity with clinical outcomes. One such signature, the 42-gene Severe-or-Mild (SoM) signature, was found to be associated with several risk factors for poor outcomes from infection, mortality, and capture the response to treatment . Similar approaches posit that deviations from health, including aging and disease, have a common systemic immune dysregulation footprint. The 150-gene signature Immune Health Metric (IHM) score represents one such effort to capture this shared influence (Sparks et al., 2024). Specifically, the IHM correlates with aging in healthy individuals and provides differential values between healthy individuals and those with monogenic diseases. In this use case, we explore TENSE and DIRES as descriptors of SoM and IHM scores. To calculate TENSE and DIRES, we assume that the SoM and IHM genes are differentially expressed, map enriched processes to DIRES, and calculate TENSE for each SoM and IHM score.
3.1.3.2 Approach and results
As described in use case 1, replicating the procedures established to calculate DIRES and TENSE, we determine the following TENSE values. Notably, as IHM is calculated as the difference between the geometric means of two gene sets (Sparks et al., 2024), (herein referred to as IHM-A and IHM-B), we calculate TENSE and DIRES for each.
For SoM: TENSESoM = 0.4 (40%).
For IHM-A: TENSEIHM-A = 0.
For IHM-B: TENSEIHM-B = 0.217 (21.7%).
For IHM, we furthermore calculate the ΔTENSE = TENSEIHM-A – TENSEIHM-B = −21.7% (Figure 10).
Figure 10
4 Discussion
This study introduces TENSE (quanTifiEd immuNe-aging dySregulation indEx), a framework that turns pathway enrichment results into a quantitative, system-level measure of immune-aging–related dysregulation. Instead of generating typical lists of significant pathways for qualitative interpretation, TENSE integrates signals from multiple pathways into key biological categories, summarizing them as an understandable dysregulation profile. As a result, it offers a higher-level overview of enrichment results, making it easier to compare biological situations based on the degree and type of aging-related regulatory imbalance present. The classification used in TENSE proved reasonably reliable, giving consistent results across repeated tests and closely matching human expert assessments. Similar outcomes were observed with alternative large language models such as Claude Haiku, suggesting that the underlying classification task is well-structured and produces stable results across different tools. Each of the five DIRES axes was found in at least one of the analyzed datasets, indicating that TENSE captures a variety of dysregulation patterns rather than being limited to just one category. Overall, these findings support TENSE as a reproducible and easy-to-interpret approach for converting complex enrichment outputs into meaningful estimates of immune-aging dysregulation.
4.1 Critical analysis of TENSE use cases
Application of the TENSE framework across transcriptional modules and experimental contrasts revealed distinct patterns of aging-associated immune dysregulation enrichment patterns across biological contexts. Within Alzheimer’s disease (AD) modules derived from WGCNA analysis, TENSE values were primarily driven by the inflammaging axis, with the strongest contributions observed in the CM8 microglial module (TENSE = 0.208/20.8%) and the brown module (TENSE = 0.131/13.1%). This pattern is consistent with extensive evidence implicating microglial activation and chronic innate immune signaling as key contributors to AD pathology. Conversely, neuronal modules such as CM1 exhibited relatively low TENSE values, with minor contributions from both inflammaging and nucleic acid sensing pathways, suggesting that immune-aging–related dysregulation is less pronounced within neuronal transcriptional programs. Similarly, the comparison between neurofibrillary tangles and plaques demonstrated moderate TENSE values largely attributable to inflammaging, a finding corroborated by the role of inflammatory signaling in AD-associated transcriptional dysregulation.
Outside the AD context, distinct enrichment patterns emerged. In the radiation response experiments, TENSE values were dominated by the DNA damage axis, particularly in the radiosensitive condition at 12 h (RS_12h; TENSE = 0.364/36.4%). This observation is consistent with the induction of genomic instability and tandem DNA damage–response pathways following irradiation, representing expected transcriptional responses to ionizing radiation (). Conversely, the sepsis-associated SOM gene signature yielded the highest TENSE value observed in the dataset (0.4/40%), driven entirely by the inflammaging axis, highlighting the extensive inflammatory activation characteristic of systemic infection. The Immune Health Metric gene signatures further illustrated this contrast: while IHM_A showed no detectable DIRES contributions, IHM_B exhibited a strong inflammaging signal (TENSE = 0.217/21.7%), suggesting that subsets of immune health–associated gene programs may capture inflammatory activation states. Notably, ΔTENSE for the two IHM signatures is negative, a finding that is in accordance with SoM and IHM scores, as expected to be inversely correlated (): DIRES annotation suggests that high IHM reflects a shift from IHM-B contributors, i.e., Inflammaging and Nucleic Acid Sensing (higher TENSE values) to IHM-A contributors, primary nucleic acid sensing mechanisms, and those associated epigenetic drift at a quiescent state (lower TENSE values).
Finally, the TM10 RNA metabolism module exhibited low but distributed contributions across multiple DIRES axes, including DNA damage, DNA repair, epigenetic drift, and inflammaging, producing a comparatively modest TENSE value. This pattern suggests diffuse engagement of multiple regulatory processes rather than dominance of a single dysregulated mechanism.
Collectively, these results illustrate how the TENSE framework can distinguish between biological contexts characterized by inflammation-dominated dysregulation (e.g., microglial activation and sepsis) and those dominated by genomic stress responses (e.g., radiation exposure). Such mechanistic differentiation is difficult to infer directly from lists of enriched pathways, highlighting the utility of summarizing pathway enrichment outputs into structured, interpretable dysregulation profiles. Table 4 summarizes DIRES and TENSE scoring across the use cases.
Table 4
| Use cases | D_rlt | I_rlt | R_rlt | E_rlt | S_rlt | TENSE |
|---|---|---|---|---|---|---|
| AD CM1 neuronal | 0 | 0.029 | 0 | 0 | 0.013 | 0.022 |
| AD CM8 microglial | 0 | 0.208 | 0 | 0 | 0 | 0.208 |
| AD module brown | 0 | 0.131 | 0 | 0 | 0 | 0.131 |
| AD module turquoise | 0.005 | 0.038 | 0 | 0 | 0.009 | 0.012 |
| AD NFT vs. plaques | 0 | 0.111 | 0 | 0 | 0.009 | 0.096 |
| AD Pink | 0 | 0 | 0 | 0 | 0 | 0 |
| AD Yellow | 0 | 0 | 0 | 0 | 0 | 0 |
| IHM_A | 0 | 0 | 0 | 0 | 0 | 0 |
| IHM_B | 0 | 0.217 | 0 | 0 | 0 | 0.217 |
| RR_12 h | 0.2 | 0 | 0 | 0 | 0 | 0.2 |
| RR_4 h | 0.2 | 0 | 0 | 0 | 0 | 0.2 |
| RS_12 h | 0.364 | 0 | 0 | 0 | 0 | 0.364 |
| RS_4 h | 0.044 | 0.044 | 0 | 0 | 0 | 0.044 |
| SOM | 0 | 0.4 | 0 | 0 | 0 | 0.4 |
| TM10 RNA metabolism | 0.044 | 0.013 | 0.033 | 0.013 | 0.002 | 0.025 |
Summary of use cases and classification by TENSE and DIRES.
Use cases are presented in alphabetical order. The “_rlt” suffix denotes the utilization of relative scores (i.e., , i ∈ D, I, R, E, S) for each DIRES category.
4.2 Limitations, caveats, and further development
In this study, the TENSE and, by extension, DIRES are defined as an alternative to arbitrary annotation of pathway enrichment data and forming decisions about the contribution and directionality of biological disorder in various settings. Rather than a comprehensive annotation of every possible process, it adopts a utilitarian approach, annotating hubs of disorder as aging progresses and providing a summary estimate that translates dysfunction as measurable relative disorder. The use cases explored indicate that TENSE and DIRES align with the biology and novel findings of the original studies, showing concordance in annotating data stemming from similar biological contexts and capturing the directionality of biological processes responding to exposures.
TENSE should be viewed within the appropriate context and limitations. First, it should be acknowledged that aging is a complicated process and even more so at the molecular level. The pathways selected to construct the TENSE workflow are hypothesis-driven rather than data-driven themselves. This issue is addressed by selecting hub processes that reflect molecular events that are not specifically modeled. For example, mtDNA leakage in the cytoplasm may not be directly modeled as a term in the formula; however, it serves as a proxy in the Nucleic Acid Sensing term. Similarly, genotoxic stress from various agents or metabolic/proteostatic effects is not directly modeled by a dedicated term but would correspondingly score as a proxy in the D, R, and I categories – as shown in the relevant use case 2.
Another caveat is that the theoretical formulation of TENSE involves inherent skew. For instance, the TENSE formulation processes D, I, R, E, S as additive and equivalent contributors. This potential caveat is addressed by substituting each constituent with its standardized equivalent and finally standardizing the output to its maximum value. This, in turn, yields interpretable values: higher TENSE values reflect greater relative utilization of the transcriptome, with DIRES revealing the relative contribution of each component. LLM stochasticity may produce TENSE values that are generally reliable based on our own iterative validation; however, the human consensus annotation herein is performed by the authors for proof of concept. A broader concept that utilizes multiple raters, e.g., a Delphi consensus and reporting inter-rater agreement statistics, can further bolster human-reference reliability and provide a more direct comparator to LLM stochasticity.
Similarly, it is worth reiterating, as a potential caveat, that TENSE does not capture the detailed contributions of each possible contributor to aging, but rather summarizes several salient hub processes. However, this limitation can be further overcome by adding additional hub functions in the DIRES schema, e.g., mitochondrial function, and adjusting the TENSE formulation correspondingly. As a modular workflow, TENSE relies on a semantic layer that can be adapted to a specific concept. Rather than simply representing the weighted standardized sum of activity scores and imposing the arithmetic mean as a constraint for each DIRES term, it aims to reduce skew by biologically similar overrepresented processes. We have opted to constrain potential scoring inflation of DIRES terms by infection terms that report on redundant gene networks, such as ribosomal proteins. Hence, these pathways do not score explicitly in DIRES unless referring to specific inflammatory processes during the course of infection.
Another potential caveat of the TENSE framework is that its initial validation was restricted to a small selection of reference conditions, specifically one interferon-associated module and one housekeeping module. Although this demonstrates that the method can work, it does not confirm its applicability to a wide range of biological scenarios. To address this issue, future studies should include a more diverse set of baseline and homeostatic conditions, enabling a thorough evaluation of the framework’s robustness and accuracy, especially in situations close to floor effects or lower TENSE states. The TENSE workflow should also be considered within the context and constraints of overrepresentation analyses, in that they represent enrichment or depletion of gene sets against our knowledge base, rather than in vivo evidence of pathway activity. As such, high-throughput datasets containing, e.g., methylomics, differential gene expression, and DNA damage marker data can provide alternative formulations of TENSE and benchmarking closer to the theoretical model, rather than the current enrichment pattern-based formulation.
Lastly, possible permutations that utilize the same theoretical model are possible. For instance, another pathway-centric approach could incorporate Shannon Entropy (Shannon, 1948), which estimates pathways related to aging and standardizes their impact over the background in sequencing experiments (; Zenil et al., 2018). Furthermore, it should be noted that it was not conceptualized as a substitute or surrogate to existing predictors of biological age. Following Horvath’s clock (), several approaches utilizing data from epigenomics (Porter et al., 2021, , Prosz et al., 2024), metabolomics (), transcriptomics (Zakar-Polyák et al., 2024) and proteomics () based approaches, among others () have provided deeper insight into deriving accurate biological age estimates (). By design, TENSE and DIRES were constructed to capture a distinct biological signal from aging clocks. While a biological clock would be utilized to estimate biological age, TENSE would be utilized to estimate the effect of an exposure that maps to age-related processes, or the skew of any of these processes (e.g., inflammation and/or DNA damage) in a biological context (e.g., disease).
TENSE can thus complement these approaches, functioning more like a biological compass and accelerometer, rather than a biological clock.
5 Conclusion
The TENSE and DIRES scoring schema represented a novel, reductionist approach in annotating pathway enrichment data and an alternative to arbitrary or hypothesis-driven interpretation. The responsiveness of both TENSE and DIRES to complex disease processes, cell-specific gene modules, and complex exposures, including ionizing radiation, across diverse biological contexts, suggests further potential usage beyond aging-related disease. Using the provided cases, we highlighted the potential of TENSE as a reproducible and interpretable method for summarizing complex pathway enrichment outputs into biologically meaningful estimates of immune-aging–related dysregulation. In this sense, TENSE transforms pathway enrichment outputs from descriptive lists into quantitative, mechanistically interpretable estimates of system-level dysregulation.
Statements
Data availability statement
The original contributions presented in the study are included in the article/Supplementary material; further inquiries can be directed to the corresponding author/s.
Ethics statement
Ethical approval was not required for the study involving humans in accordance with the local legislation and institutional requirements. Written informed consent to participate in this study was not required from participants or their legal guardians/next of kin in accordance with national legislation and institutional requirements.
Author contributions
GV: Conceptualization, Data curation, Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing. GH: Funding acquisition, Writing – original draft, Writing – review & editing.
Funding
The author(s) declared that financial support was not received for this work and/or its publication.
Acknowledgments
The authors would like to express their gratitude to the Cyprus Academy of Sciences, Letters, and Arts for their support during the preparation of this manuscript. GV would like to express his gratitude to Taisia Kyriakou for her insight and support during the preparation of this manuscript. Figures were created through BioRender.
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
The author GV declared that they were an editorial board member of Frontiers, at the time of submission. This had no impact on the peer review process and the final decision.
Generative AI statement
The author(s) declared that Generative AI was not used in the creation of this manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Supplementary material
The Supplementary material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/frai.2026.1732901/full#supplementary-material
Footnotes
1.^KNIME Workflow version 5.10; available from https://www.knime.com; Accessed March 9th 2026
2.^Available from Raw data usage from referenced studies, Enrichr access to Reactome Library version 2024; Available from https://maayanlab.cloud/Enrichr/enrich; Accessed March 9th 2026
3.^https://openai.com/index/gpt-4-1
4.^Available from: https://developers.openai.com/api/docs/models/gpt-4.1-nano; Accessed 09/03/2026
5.^Library version 2024; Available from https://maayanlab.cloud/Enrichr/enrich
6.^version 5.10; available from https://www.knime.com; run locally
References
1
ArgentieriM. A.XiaoS.BennettD.WinchesterL.Nevado-HolgadoA. J.GhoseU.et al. (2024). Proteomic aging clock predicts mortality and risk of common age-related diseases in diverse populations. Nat. Med.30, 2450–2460. doi: 10.1038/s41591-024-03164-7,
2
AzamM.ChenY.ArowoloM. O.LiuH.PopescuM.XuD. (2024). A comprehensive evaluation of large language models in mining gene relations and pathway knowledge. Quant Biol12, 360–374. doi: 10.1002/qub2.57,
3
Bertucci-RichterE. M.ParrottB. B. (2023). The rate of epigenetic drift scales with maximum lifespan across mammals. Nat. Commun.14:7731. doi: 10.1038/s41467-023-43417-6,
4
Bertucci-RichterE. M.ShealyE. P.ParrottB. B. (2024). Epigenetic drift underlies epigenetic clock signals, but displays distinct responses to lifespan interventions, development, and cellular dedifferentiation. Aging (Albany NY)16, 1002–1020. doi: 10.18632/aging.205503,
5
Borràs-FresnedaM.BarquineroJ. F.GomolkaM.HornhardtS.RösslerU.ArmengolG.et al. (2016). Differences in DNA repair capacity, cell death and transcriptional response after irradiation between a radiosensitive and a Radioresistant cell line. Sci. Rep.6:27043. doi: 10.1038/srep27043,
6
CaoW. (2022). IFN-aging: coupling aging with interferon response. Front. Aging3:870489. doi: 10.3389/fragi.2022.870489,
7
CarelsN. (2024). Assessing RNA-Seq workflow methodologies using Shannon entropy. Biology (Basel)13:482. doi: 10.3390/biology13070482,
8
CarstensenB. (2010). Repeatability, Reproducibility and Coefficient of Variation, Comparing Clinical Measurement Methods. Hoboken, NJ: Wiely, 107–114.
9
ChenS.ChangY.LiL.AcostaD.LiY.GuoQ.et al. (2022). Spatially resolved transcriptomics reveals genes associated with the vulnerability of middle temporal gyrus in Alzheimer's disease. Acta Neuropathol. Commun.10:188. doi: 10.1186/s40478-022-01494-6,
10
ChenJ.ChenY.LuY.HeY.JiangF.HuL.et al. (2026). Spatial and multi-omics transcriptomic dissects platinum resistance in lung adenocarcinoma: a five-gene predictive model with tumor microenvironment dynamics. Chem. Biol. Interact.429:111952. doi: 10.1016/j.cbi.2026.111952,
11
ChenE. Y.TanC. M.KouY.DuanQ.WangZ.MeirellesG. V.et al. (2013). Enrichr: interactive and collaborative HTML5 gene list enrichment analysis tool. BMC Bioinformatics14:128. doi: 10.1186/1471-2105-14-128,
12
CoppM. E.ShineJ.BrownH. L.NimmalaK. R.HansenO. B.ChubinskayaS.et al. (2023). Sirtuin 6 activation rescues the age-related decline in DNA damage repair in primary human chondrocytes. Aging (Albany NY)15, 13628–13645. doi: 10.18632/aging.205394,
13
DasS.LiZ.WachterA.AllaS.NooriA.AbdourahmanA.et al. (2024). Distinct transcriptomic responses to aβ plaques, neurofibrillary tangles, and APOE in Alzheimer's disease. Alzheimers Dement.20, 74–90. doi: 10.1002/alz.13387,
14
DorrityT. J.ShinH.GertieJ. A.ChungH. (2024). The sixth sense: self-nucleic acid sensing in the brain. Adv. Immunol.161, 53–83. doi: 10.1016/bs.ai.2024.03.001,
15
FranceschiC.SalvioliS.GaragnaniP.De EguileorM.MontiD.CapriM. (2017). Immunobiography and the heterogeneity of immune responses in the elderly: a focus on Inflammaging and trained immunity. Front. Immunol.8:982. doi: 10.3389/fimmu.2017.00982,
16
FulopT.LarbiA.PawelecG.KhalilA.CohenA. A.HirokawaK.et al. (2023). Immunology of aging: the birth of Inflammaging. Clin. Rev. Allergy Immunol.64, 109–122. doi: 10.1007/s12016-021-08899-6,
17
FülöpT.LarbiA.WitkowskiJ. M. (2019). Human Inflammaging. Gerontology65, 495–504. doi: 10.1159/000497375,
18
GalkinF.MamoshinaP.AliperA.PutinE.MoskalevV.GladyshevV. N.et al. (2020). Human gut microbiome aging clock based on taxonomic profiling and deep learning. iScience23:101199. doi: 10.1016/j.isci.2020.101199,
19
GanesanA.MooreA. R.ZhengH.TohJ.FreedmanM.MagisA. T.et al. (2025). A conserved immune dysregulation signature is associated with infection severity, risk factors prior to infection, and treatment response. Immunity58, 2104–2119.e5. doi: 10.1016/j.immuni.2025.05.020,
20
GliechC. R.HollandA. J. (2020). Keeping track of time: the fundamentals of cellular clocks. J. Cell Biol.219. doi: 10.1083/jcb.202005136,
21
GorbunovaV.SeluanovA.MaoZ.HineC. (2007). Changes in DNA repair during aging. Nucleic Acids Res.35, 7466–7474. doi: 10.1093/nar/gkm756,
22
HanifF.RazaA.MohammedH. A. (2026). Compact deep learning models for colon histopathology focusing performance and generalization challenges. Sci. Rep.16:5489. doi: 10.1038/s41598-026-35119-y,
23
HannumG.GuinneyJ.ZhaoL.ZhangL.HughesG.SaddaS.et al. (2013). Genome-wide methylation profiles reveal quantitative views of human aging rates. Mol. Cell49, 359–367. doi: 10.1016/j.molcel.2012.10.016,
24
HatchE. M.HetzerM. W. (2015). Linking micronuclei to chromosome fragmentation. Cell161, 1502–1504. doi: 10.1016/j.cell.2015.06.005,
25
HongE. K.LeeS.SongO.LeungA.HammerM.SuhC. H. (2025). Temperature setting of a multimodal generative artificial intelligence (AI) model: association with accuracy and quality of AI-generated chest radiograph reports. AJR Am. J. Roentgenol.226:e2533704. doi: 10.2214/AJR.25.33704,
26
HorvathS. (2013). DNA methylation age of human tissues and cell types. Genome Biol.14:R115. doi: 10.1186/gb-2013-14-10-r115,
27
HuangH.ChenY.XuW.CaoL.QianK.BischofE.et al. (2025). Decoding aging clocks: new insights from metabolomics. Cell Metab.37, 34–58. doi: 10.1016/j.cmet.2024.11.007,
28
IssaJ. P. (2014). Aging and epigenetic drift: a vicious cycle. J. Clin. Invest.124, 24–29. doi: 10.1172/JCI69735,
29
JansenR.HanL. K. M.VerhoevenJ. E.AbergK. A.van den OordE. C. G. J.MilaneschiY.et al. (2021). An integrative study of five biological clocks in somatic and mental health. eLife10:e59479. doi: 10.7554/eLife.59479,
30
LaffonB.BonassiS.CostaS.ValdiglesiasV. (2021). Genomic instability as a main driving factor of unsuccessful ageing: potential for translating the use of micronuclei into clinical practice. Mutat. Res. Rev. Mutat. Res.787:108359. doi: 10.1016/j.mrrev.2020.108359,
31
LemusA. J. J.TeweldeE.BhalaR.GuanA.XuA.BenayounB. A. (2026). The interplay of epigenetic remodelling and transposon-mediated genomic instability in ageing and longevity. Open Biol.16:250093. doi: 10.1098/rsob.250093,
32
LestienneR. (1988). From physical to biological time. Mech. Ageing Dev.43, 189–228. doi: 10.1016/0047-6374(88)90032-2,
33
LiuL.GaoM.HeW.WangM.ZhouS.WangX.et al. (2026). Identification of mitochondrial permeability transition-related gene signatures to predict lung adenocarcinoma survival and drug response. Cell Cycle25, 1–35. doi: 10.1080/15384101.2025.2606113,
34
López-GilL.Pascual-AhuirA.ProftM. (2023). Genomic instability and epigenetic changes during aging. Int. J. Mol. Sci.24:279. doi: 10.3390/ijms241814279,
35
López-OtínC.BlascoM. A.PartridgeL.SerranoM.KroemerG. (2023). Hallmarks of aging: an expanding universe. Cell186, 243–278. doi: 10.1016/j.cell.2022.11.001,
36
MaZ.WangH.CaiY.WangH.NiuK.WuX.et al. (2018). Epigenetic drift of H3K27me3 in aging links glycolysis to healthy longevity in Drosophila. eLife7:e35368. doi: 10.7554/eLife.35368,
37
MillerK. N.VictorelliS. G.SalmonowiczH.DasguptaN.LiuT.PassosJ. F.et al. (2021). Cytoplasmic DNA: sources, sensing, and role in aging and disease. Cell184, 5506–5526. doi: 10.1016/j.cell.2021.09.034,
38
MiyakeT.ShimadaM.MatsumotoY.OkinoA. (2019). DNA damage response after ionizing radiation exposure in skin keratinocytes derived from human-induced pluripotent stem cells. Int. J. Radiat. Oncol. Biol. Phys.105, 193–205. doi: 10.1016/j.ijrobp.2019.05.006,
39
MooreA. R.ZhengH.GanesanA.Hasin-BrumshteinY.MaddaliM. V.LevittJ. E.et al. (2025). A consensus immune dysregulation framework for sepsis and critical illnesses. Nat. Med.31, 4084–4096. doi: 10.1038/s41591-025-03956-5,
40
MorabitoS.MiyoshiE.MichaelN.SwarupV. (2020). Integrative genomics approach identifies conserved transcriptomic networks in Alzheimer's disease. Hum. Mol. Genet.29, 2899–2919. doi: 10.1093/hmg/ddaa182,
41
NguyenT. M.ShafiA.NguyenT.DraghiciS. (2019). Identifying significantly impacted pathways: a comprehensive review and assessment. Genome Biol.20:203. doi: 10.1186/s13059-019-1790-4,
42
PapatriantafyllouM. (2013). DNA damage sensor in the interferon response. Nat. Rev. Immunol.13, 155–155. doi: 10.1038/nri3418
43
PaulB. D.SnyderS. H.BohrV. A. (2021). Signaling by cGAS-STING in neurodegeneration, Neuroinflammation, and aging. Trends Neurosci.44, 83–96. doi: 10.1016/j.tins.2020.10.008,
44
PoretskyE.BlakeV. C.AndorfC. M.SenT. Z. (2025). Assessing the performance of generative artificial intelligence in retrieving information against manually curated genetic and genomic data. Database (Oxford)2025:11. doi: 10.1093/database/baaf011,
45
PorterH. L.BrownC. A.RoopnarinesinghX.GilesC. B.GeorgescuC.FreemanW. M.et al. (2021). Many chronological aging clocks can be found throughout the epigenome: implications for quantifying biological aging. Aging Cell20:e13492. doi: 10.1111/acel.13492,
46
ProszA.PipekO.BörcsökJ.PallaG.SzallasiZ.SpisakS.et al. (2024). Biologically informed deep learning for explainable epigenetic clocks. Sci. Rep.14:1306. doi: 10.1038/s41598-023-50495-5,
47
RazaA.HanifF. (2026). Chronological review and performance analysis of YOLO-based deep learning frameworks in complex geospatial environments. Appl. Artif. Intell.40:2626114. doi: 10.1080/08839514.2026.2626114
48
RazaA.HanifF.MohammedH. A. (2025a). Efficient crack and surface-type recognition via CNN-block development mechanism and edge profiling. Sci. Rep.15:40073. doi: 10.1038/s41598-025-25956-8,
49
RazaA.HanifF.MohammedH. A. (2025b). Analyzing the enhancement of CNN-YOLO and transformer based architectures for real-time animal detection in complex ecological environments. Sci. Rep.15:39142. doi: 10.1038/s41598-025-26645-2,
50
RazaA.HanifF.MohammedH. A. (2026). Clinical validation of lightweight CNN architectures for reliable multi-class classification of lung cancer using histopathological imaging techniques. Sci. Rep.16:6512. doi: 10.1038/s41598-026-36652-6,
51
SchumacherB.PothofJ.VijgJ.HoeijmakersJ. H. J. (2021). The central role of DNA damage in the ageing process. Nature592, 695–703. doi: 10.1038/s41586-021-03307-7,
52
ShannonC. E. (1948). A mathematical theory of communication. Bell Syst. Tech. J.27, 379–423. doi: 10.1002/j.1538-7305.1948.tb01338.x
53
ShustermanR.WatersA. C.O'NeillS.BangsM.LuuP.TuckerD. M. (2025). An active inference strategy for prompting reliable responses from large language models in medical practice. NPJ Digit. Med.8:119. doi: 10.1038/s41746-025-01516-2,
54
SongX.TouX.LiL.WangF.QianC.WangR.et al. (2026). Efferocytosis-associated transcriptomic patterns characterize prognosis and immune landscape in osteosarcoma. J Bone Oncol57:100743. doi: 10.1016/j.jbo.2026.100743,
55
SparksR.RachmaninoffN.LauW. W.HirschD. C.BansalN.MartinsA. J.et al. (2024). A unified metric of human immune health. Nat. Med.30, 2461–2472. doi: 10.1038/s41591-024-03092-6,
56
StillmanJ. M.KiniwaT.SchaferD. P. (2024). Nucleic acid sensing in the central nervous system: implications for neural circuit development, function, and degeneration. Immunol. Rev.327, 71–82. doi: 10.1111/imr.13420,
57
TaffoniC.SteerA.MarinesJ.ChammaH.VilaI. K.LaguetteN. (2021). Nucleic acid immunity and DNA damage response: new friends and old foes. Front. Immunol.12:660560. doi: 10.3389/fimmu.2021.660560,
58
TatulloM. (2024). Entropy meets physiology: should we translate aging as disorder?Stem Cells42, 91–97. doi: 10.1093/stmcls/sxad084,
59
TenchovR.SassoJ. M.WangX.ZhouQ. A. (2024). Aging hallmarks and progression and age-related diseases: a landscape view of research advancement. ACS Chem. Neurosci.15, 1–30. doi: 10.1021/acschemneuro.3c00531,
60
TiwariK.MatthewsL.MayB.ShamovskyV.Orlic-MilacicM.RothfelsK.et al. (2023). ChatGPT-4.1-NANO usage in the Reactome curation process. bioRxiv2023:195. doi: 10.1101/2023.11.08.566195,
61
VavougiosG. D.MavridisT.ArtemiadisA.KrogfeltK. A.HadjigeorgiouG. (2022). Trained immunity in viral infections, Alzheimer's disease and multiple sclerosis: a convergence in type I interferon signalling and IFNβ-1a. Biochim. Biophys. Acta Mol. basis Dis.1868:166430. doi: 10.1016/j.bbadis.2022.166430,
62
VlachogiannisN. I.NtourosP. A.PappaM.KravvaritiE.KostakiE. G.FragoulisG. E.et al. (2023). Chronological age and DNA damage accumulation in blood mononuclear cells: a linear Association in Healthy Humans after 50 years of age. Int. J. Mol. Sci.24:148. doi: 10.3390/ijms24087148,
63
WangH.DaiH.WangY.WuQ.ZhuM.YinW.et al. (2026). Leveraging a chromosomal instability-based signature to predict the prognosis and immune landscape of breast cancer. Genes Dis13:101924. doi: 10.1016/j.gendis.2025.101924,
64
WangZ.JinQ.WeiC.-H.TianS.LaiP.-T.ZhuQ.et al. (2025). GeneAgent: self-verification language agent for gene-set analysis using domain databases. Nat. Methods22, 1677–1685. doi: 10.1038/s41592-025-02748-6,
65
WhittemoreK.Martínez-NevadoE.BlascoM. A. (2019). Slower rates of accumulation of DNA damage in leukocytes correlate with longer lifespans across several species of birds and mammals. Aging (Albany NY)11, 9829–9845. doi: 10.18632/aging.102430,
66
WilmothD. R. (2012). The relationships between common measures of glucose meter performance. J. Diabetes Sci. Technol.6, 1087–1093. doi: 10.1177/193229681200600512,
67
WuG.MatthewsL.BoyerN.MilacicM.BeaversD.LiN. T.et al. (2025). Application of large language models for annotating genes into Reactome pathways. bioRxiv2025:723. doi: 10.64898/2025.12.20.695723,
68
XieZ.BaileyA.KuleshovM. V.ClarkeD. J. B.EvangelistaJ. E.JenkinsS. L.et al. (2021). Gene set knowledge discovery with Enrichr. Curr Protoc1:e90. doi: 10.1002/cpz1.90,
69
YahyapourR.AminiP.RezapourS.ChekiM.RezaeyanA.FarhoodB.et al. (2018). Radiation-induced inflammation and autoimmune diseases. Mil. Med. Res.5:9. doi: 10.1186/s40779-018-0156-7,
70
Zakar-PolyákE.CsordasA.PálovicsR.KerepesiC. (2024). Profiling the transcriptomic age of single-cells in humans. Commun Biol7:1397. doi: 10.1038/s42003-024-07094-5,
71
ZenilH.KianiN. A.TegnérJ. (2018). A review of graph and network complexity from an algorithmic information perspective. Entropy (Basel)20:551. doi: 10.3390/e20080551,
72
ZhaoY.SimonM.SeluanovA.GorbunovaV. (2023). DNA damage and repair in age-related inflammation. Nat. Rev. Immunol.23, 75–89. doi: 10.1038/s41577-022-00751-y,
73
ZhouY.ZhouB.PacheL.ChangM.KhodabakhshiA. H.TanaseichukO.et al. (2019). Metascape provides a biologist-oriented resource for the analysis of systems-level datasets. Nat. Commun.10:1523. doi: 10.1038/s41467-019-09234-6,
Summary
Keywords
Alzheimer’s disease, artificial intelligence, differential gene expression, gene expression data, inflammation, KNIME analytics platform, large language model, scoring–algorithm
Citation
Vavougios GD and Hadjigeorgiou G (2026) The quantified immune-aging dysregulation index: a large-language model-powered method for annotating and quantifying systems-level dysregulation. Front. Artif. Intell. 9:1732901. doi: 10.3389/frai.2026.1732901
Received
26 October 2025
Revised
09 May 2026
Accepted
18 May 2026
Published
08 June 2026
Volume
9 - 2026
Edited by
Jose Laffita Mesa, Karolinska Institutet (KI), Sweden
Reviewed by
Yizhi Wang, Virginia Tech, United States
Saad Ahmed, Prince Sattam Bin Abdulaziz University, Saudi Arabia
Updates
Copyright
© 2026 Vavougios and Hadjigeorgiou.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: George D. Vavougios, dantevavougios@hotmail.com; vavougyios.georgios@ucy.ac.cy
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.