Abstract
Background:
Speech and language impairments, long recognized as early symptoms of Alzheimer’s disease (AD), can now be quantified with unprecedented precision due to recent advances in natural language processing (NLP) and artificial intelligence (AI). Despite growing interest in AI-enabled speech biomarkers, few studies have linked spontaneous speech to biologically verified AD, and most have focused on English-language data or acoustic features with limited linguistic interpretability. Here, we present the first end-to-end machine-learning framework for automatic AD detection from German speech, using clinical-biological criteria validated by cerebrospinal fluid (CSF) biomarkers.
Methods:
44 participants were included: 22 biomarker-defined AD cases from a prospective observational study (German Clinical Trials Register, DRKS00030633) and 22 socio-demographically matched cognitively healthy controls (CHC). Connected speech was elicited using the standardized Cookie Theft picture description task. Recordings were transcribed with a state-of-the-art automatic speech recognition (ASR) system. From these transcripts, 32 theory-driven linguistic biomarkers were computed with an advanced NLP tool, falling into three categories: information-theoretic, lexical richness, and syntactic. AD-versus-CHC classification used five supervised models (logistic regression, support vector machine with a radial basis function kernel, random forest, gradient boosting, XGBoost) under stratified five-fold cross-validation, with stability-based recursive feature elimination performed within training folds. Model interpretability was assessed using SHapley Additive exPlanations (SHAP).
Results:
Recursive feature elimination retained seven of 32 candidate speech biomarkers as consistently informative across folds. Trained on this subset, all classifiers showed strong discrimination between biomarker-defined AD and CHC. Logistic regression, SVM, random forest, and gradient boosting achieved ~91% mean accuracy with F1 ≈ 0.90 and sensitivity ≈ 0.90, while XGBoost was slightly lower (~89% accuracy). SHAP analyses indicated that model decisions were primarily driven by information-theoretic and structural markers: lower compressibility, reduced lexical density, shorter clauses and sentences, and weaker predictive sequencing indexed by higher-order n-gram statistics.
Conclusion:
Clinically meaningful linguistic biomarkers can be robustly derived from spontaneous speech, even in small, well-characterized clinical samples. Theory-driven features and stability-focused modeling show that information-theoretic and structural properties of connected speech capture core Alzheimer-related impairments with robust classification performance. These findings support AI-enabled speech analysis as a non-invasive, scalable complement to established biological biomarkers of Alzheimer’s disease.
1 Introduction
The rapid growth of the aging population has intensified global concerns over what is increasingly described as a “dementia epidemic.” The prevalence of dementia has risen sharply, currently affecting an estimated 57 million people worldwide (1) and is projected to triple by 2050 (2). This surge poses a major public health challenge, profoundly affecting individuals’ well-being while placing an immense strain on healthcare systems and social infrastructure worldwide. The estimated annual global cost of dementia exceeds US$1.3 trillion annually and is expected to rise to US$2.8 trillion by 2030, outpacing expenditures on cancer and cardiovascular disease (3).
Alzheimer’s disease (AD)–the most common cause of dementia–is a progressive neurodegenerative disorder characterized by cognitive decline, memory loss, and behavioral changes (4). Its neuropathological hallmarks include the accumulation of amyloid-β (Aβ) plaques and tau neurofibrillary tangles in the brain, processes that may begin decades before the onset of clinical symptoms (5, 6). Early, accurate diagnosis is therefore essential for effective disease management and timely therapeutic intervention, as it provides a critical window for early pharmacological and non-pharmacological interventions that may slow neurodegenerative processes. In this context, the first disease-modifying treatments have been approved, namely Lecanemab, a humanized immunoglobulin G1 monoclonal antibody targeting soluble amyloid-beta protofibrils (7), and Donanemab, an immunoglobulin G1 monoclonal antibody directed against N-terminal pyroglutamate-modified amyloid-beta (8).
The gold standard for diagnosing AD relies on the integration of multiple complementary methods, including neuroimaging, cerebrospinal fluid (CSF) biomarker analysis, and comprehensive clinical and neuropsychological assessment. This integrative framework reflects the International Working Group (IWG) model, which defines AD through the convergence of biological markers and clinical phenotype (9–12). Under the IWG criteria, diagnostic accuracy arises from aligning biomarker evidence with cognitive and clinical presentation, enabling a more precise characterization of molecular pathology, neural dysfunction, and cognitive impairment.
However, each of these methods carries practical and methodological limitations that restrict their clinical scalability and widespread implementation. Among these, neuroimaging techniques provide in vivo visualization of the structural and molecular substrates of AD (13). Magnetic resonance imaging (MRI) characterizes macroscopic structural alterations—including regional atrophy, cortical thinning, and white-matter hyperintensities—reflecting neuronal loss and cerebrovascular burden (14). Positron emission tomography (PET) complements this by detecting amyloid-β (Aβ) and tau aggregates using specific radioligands, thereby mapping the core molecular pathologies that define the disease (6, 15). Aβ plaques accumulate extracellularly in neocortical regions as an early event in pathogenesis (4, 16), whereas tau neurofibrillary tangles develop intraneuronally, originating in the medial temporal lobe and spreading to association cortices in parallel with neurodegeneration and cognitive decline (17, 18). Although invaluable for characterizing AD pathology, neuroimaging methods remain expensive, time-consuming, and technically demanding, requiring specialized infrastructure and expert interpretation. PET, in particular, is further constrained by radiotracer availability, radiation exposure, and limited clinical accessibility (19). CSF biomarkers provide a biochemical index of AD pathology (20, 21). Decreased Aβ₁–₄₂ and elevated phosphorylated and total tau concentrations reflect amyloid and tau pathology and correlate strongly with PET imaging (22). The Aβ₁–₄₂/Aβ₁–₄₀ ratio further improves diagnostic specificity by accounting for interindividual variability in Aβ production (23, 24). Despite its high analytical validity, lumbar puncture remains invasive, limiting patient acceptability and scalability (25). Blood-based biomarkers, including plasma Aβ₄₂/₄₀ and p-tau217, show strong correspondence with CSF and PET measures but remain subject to pre-analytical variability, peripheral metabolism, and incomplete assay standardization (26–29).
A critical limitation shared by current imaging and fluid biomarkers is that, although they sensitively index Alzheimer’s-related pathology and show meaningful correlations with cognitive and clinical measures, these biological markers alone cannot fully establish the presence or severity of cognitive impairment, nor can they directly capture functional disease progression, which is essential for quantifying treatment response in interventional trials. This limitation is underscored by longitudinal studies of biomarker-positive but cognitively unimpaired individuals: even markedly abnormal amyloid and tau profiles confer only probabilistic risk, with a substantial proportion of such individuals remaining clinically stable rather than progressing to symptomatic AD over typical follow-up periods, highlighting persistent uncertainty at the level of individual diagnosis (30, 31).
Neuropsychological assessments remain essential for quantifying the cognitive and behavioral manifestations of Alzheimer’s disease. Comprehensive batteries such as the Committee to Establish a Registry for Alzheimer’s Disease Neuropsychological Assessment Battery (CERAD-NAB) (32)—which integrates episodic memory, language, and praxis tasks into standardized composite scores—together with widely used global screening instruments, including the Montreal Cognitive Assessment (MoCA) (33) and the Mini-Mental State Examination (MMSE) (34), provide standardized indices of cognitive status and longitudinal change. However, many of these instruments were originally developed to detect manifest dementia rather than subtle early-stage impairment, limiting their sensitivity in prodromal or preclinical phases. In addition, their ecological validity is restricted (35), they are susceptible to floor and ceiling effects (36), and repeated administration induces practice effects (37), collectively reducing psychometric robustness in early disease stages (36).
Despite advances in the early detection of Alzheimer’s disease (AD) using fluid and imaging biomarkers, there remains a clear need for complementary measures that are non-invasive, cost-effective, and scalable. In this context, digital biomarkers have been proposed as valuable adjuncts to biological and neuropsychological assessments, providing ecologically valid indicators of cognitive functioning in everyday settings (38). According to the European Medicines Agency (EMA), a digital biomarker is “an objective, quantifiable measure of physiology and/or behavior used as an indicator of a biological or pathological process, or a response to an exposure or an intervention, that is derived from a digital measure” (39, p. 4). This definition explicitly incorporates behavioral signals alongside physiological ones, reinforcing the relevance of functional measures in disease characterization. This perspective aligns with the broader transition toward digital medicine, where advances in computational phenotyping—together with the widespread availability of smartphones, tablets, wearables, and sensor-enabled devices—enable data-driven characterization of cognitive functioning beyond the clinical setting (40–42).
Within this evolving landscape, speech and language measures constitute a class of digital behavioral biomarkers, long recognized as sensitive indicators of cognitive functioning, as evidenced by their disruption across diverse neurodegenerative conditions (43, 44). Rather than indexing molecular pathology directly, speech captures the functional expression of underlying neurobiological changes in memory, executive control, and distributed language networks. Recent developments in natural language processing (NLP) and artificial intelligence (AI) have substantially enhanced the precision, scalability, and interpretability with which such speech-based digital biomarkers can be quantified, creating unprecedented opportunities to detect subtle linguistic markers relevant to early cognitive impairment and Alzheimer’s disease (45–47).
Several studies have therefore focused specifically on spontaneous (connected) speech, as it most closely approximates how individuals communicate in everyday life and imposes minimal task-specific constraints. Unlike structured naming, repetition, or list-learning paradigms, spontaneous speech requires speakers to generate and sustain a self-directed communicative goal and to organize content without external scaffolding. As a result, spontaneous speech production places substantial demands on online working-memory resources for syntactic formulation, semantic memory access, and the coordinated engagement of distributed language-production networks, as speakers must plan, formulate, and monitor utterances in real time (48–51).
Early prospective evidence linking spontaneous speech to preclinical AD came from Cuetos et al. (52), who compared connected-speech samples from asymptomatic Presenilin-1 mutation carriers and non-carriers and found markedly reduced semantic content in carriers decades before expected symptom onset. Retrospective analyses further corroborate this relationship. Linguistic simplification, reduced syntactic complexity, and impoverished vocabulary were evident in Iris Murdoch’s unedited published novels in the years preceding her clinical diagnosis (53). Longitudinal analyses of Ronald Reagan’s unscripted public speech similarly revealed increasing filled pauses and declining lexical diversity well before his documented cognitive impairment (54). Neuropathologically confirmed evidence reinforces this pattern: Ahmed et al. (55) demonstrated progressive declines in semantic informativeness and discourse efficiency across the transition from mild cognitive impairment to mild and moderate Alzheimer’s dementia.
The most comprehensive and methodologically rigorous validation of the associations between language, cognitive functioning, and dementia comes from the Nun Study—widely regarded as one of the most influential longitudinal aging cohorts. Snowdon et al. (56) first showed that linguistic expression in early-adult autobiographies, quantified through idea density and syntactic structure, predicts both late-life cognitive impairment and Alzheimer’s neuropathology. This association was subsequently replicated and extended across the full Nun Study cohort—established between 1991 and 1993 and comprising 678 nuns aged 75–102, with approximately 600 autopsies completed—which has generated uniquely high-quality data driving major insights into aging and dementia for more than three decades. Owing to the cohort’s remarkable homogeneity and its systematically structured methodology—including epidemiological data (autobiographies, academic records, medical histories), longitudinal assessments (daily-life activities and cognitive testing), biological sample collection (blood and genetic data), and comprehensive neuropathology (gross pathology, histology, and digital pathology)—the Nun Study offers an exceptionally consistent empirical framework for linking linguistic, cognitive, and biological measures across the lifespan. This body of work—now comprehensively synthesized in Clarke et al. (57)—consistently demonstrates that individuals exhibiting richer linguistic profiles in young adulthood show markedly reduced vulnerability to cognitive decline, whereas lower linguistic complexity predicts susceptibility across both clinical and neuropathological outcomes. These converging findings demonstrate that spontaneous (connected) speech captures core cognitive and linguistic vulnerabilities long before clinical diagnosis and indexes functional dimensions of Alzheimer’s disease that cannot be inferred from biological pathology alone.
Translating AI-enabled digital speech and language biomarkers into clinical practice demands rigorous standardization of speech-elicitation paradigms and robust frameworks for identifying clinically meaningful, theory-grounded, and human-interpretable linguistic features—core requirements for responsible AI and real-world deployment (58). Review papers have emphasized that progress in speech-based detection of cognitive impairment is fundamentally constrained by heterogeneity in elicitation tasks, recording conditions, and sample composition (59–61). Standardization of these components is critical: uncontrolled variation in linguistic demands, acoustic context, or demographic composition introduces variance that obscures disease-related signal and undermines reproducibility. A major step toward resolving these limitations was the development of the ADReSS (Alzheimer’s Dementia Recognition through Spontaneous Speech) (62) and ADReSSo (Alzheimer’s Dementia Recognition through Spontaneous Speech Only) (63) benchmark challenges, which are based on a carefully selected subset of the Pitt Corpus (64) and employ the standardized Cookie Theft picture-description task, introduced in Section 2.2. The ADReSS challenge improved this resource by releasing a carefully matched subset of the Pitt data in which Alzheimer’s and control participants were explicitly balanced for age, sex, and recording quality, thereby reducing demographic and acoustic confounds that previously limited reproducibility. The subsequent ADReSSo challenge retained the standardized elicitation task but introduced a crucial change: it released only raw speech recordings—without transcripts—thereby promoting the development and comparative evaluation of end-to-end frameworks and machine learning models for the automatic detection of AD from spontaneous speech, unlike the initial challenge, which also provided manual transcripts.
Beyond standardized elicitation and demographic control, effective clinical translation of speech-based approaches requires determining which aspects of spoken language constitute meaningful candidates for Alzheimer’s disease (AD) biomarkers. Baseline results from the ADReSSo challenge provide a clear direction: Luz et al. (63) demonstrated that models trained on linguistic features consistently outperform acoustic-only approaches, reinforcing the primacy of language in cognitive profiling. Their classifiers achieved markedly higher accuracy using linguistic features (≈77.46%) compared with acoustic models (≈64.79%), a finding echoed across independent studies (65, 66). This superiority is not only empirical but conceptual. Prior research on acoustic parameters—such as fundamental frequency, jitter, shimmer, or mel-frequency representations—has produced inconclusive and often nonspecific evidence for their diagnostic value in AD (67, 68), and these measures offer limited interpretability in clinical contexts. By contrast, measures of lexical diversity, density and sophistication, and morpho-syntactic complexity, more directly reflect the working-memory constraints and executive-function impairments that characterize the clinical presentation of Alzheimer’s disease. Linguistic features therefore represent the most reliable, interpretable, and theoretically grounded candidates for speech-based biomarkers of AD.
In parallel, a number of studies have applied deep-learning approaches to the ADReSSo challenge, using transformer-based text encoders such as BERT (69) and large language models such as LLaMA 3 8B Instruct (70), as well as acoustic encoders such as Wav2Vec 2.0 (71), to derive high-dimensional feature representations from spontaneous speech. Although these methods demonstrate the feasibility of end-to-end representation learning, their latent feature spaces are inherently opaque and provide limited interpretive value for high-stakes clinical use cases. For this reason, the present work focuses on interpretable, theory-driven markers, which support transparent, individually validated measures essential for responsible AI and clinical translation (58). Readers interested in deep-learning applications to the ADReSSo dataset are referred to Qiao et al. (72), Zhu et al. (73), Zhu et al. (74), and Shao et al. (75).
Despite extensive work on speech-based detection of AD, several methodological limitations constrain current evidence and impede clinical translation. First, the majority of studies employing the standardized Cookie Theft picture-description task have been conducted in English-speaking cohorts. Cross-linguistic investigations remain scarce, with only a few isolated efforts in Chinese (48) and Slovak (47) populations. As a result, most published findings and model behaviors are derived from English-specific linguistic structures, constraining conclusions about generalizability across languages with different morphological, syntactic, and information-structural properties (76). Second, even within existing non-English datasets, rigorous demographic control is uncommon. Many prior corpora do not systematically match participants on age, sex, and—critically—education, despite well-established effects of these variables on language production and cognitive performance. This lack of sociodemographic balancing introduces confounding variance that can artificially inflate or obscure diagnostic signal. Third, and of particular clinical relevance, most spontaneous-speech resources lack information on biological measures of pathology. As such, the predominance of clinically defined rather than biomarker-confirmed AD reduces specificity of current evidence. To our knowledge, no prior study has provided a standardized, sociodemographically matched German spontaneous-speech dataset of participants with clinical-biologically defined AD.
The present study addresses these limitations by introducing what is, to the best of our knowledge, the first end-to-end machine-learning framework for automatic AD detection from German speech in a sample of patients fulfilling clinical-biological criteria for AD and a sociodemographically matched control sample. Specifically, our aims were to:
Develop supervised AI models trained on interpretable, theory-driven feature sets extracted from standardized picture-description speech, comprising syntactic complexity, lexical density and diversity, lexical sophistication, and information-theoretic measures.
Apply stability-based recursive feature elimination within cross-validation to identify a robust subset of features consistently associated with biomarker-defined AD.
Use SHAP-based model interpretability to determine which features contribute most reliably to end-to-end AD detection, enabling transparent evaluation of the specific dimensions of connected speech that differentiate AD from controls.
2 Materials and methods
2.1 Participants and clinical measures
The AD sample (N = 22) was drawn from baseline assessments of a prospective longitudinal observational study within the Center for Dementia and Prevention in Aachen (ZDPA) at the Department of Neurology at UKA. The study was approved by the ethics committee of the Faculty of Medicine of the RWTH Aachen University (EK 384/20) and is registered in the German Clinical Trials Register (DRKS00030633). All patients provided written informed consent.
AD diagnosis followed the national guidelines for dementia diagnosis and management (S3 Leitlinien) according to the IWG criteria (9–11). Biological criteria were defined using cerebrospinal fluid concentrations of Aβ₁–₄₂, Aβ₁–₄₀, Aβ₁–₄₂/Aβ₁–₄₀ ratio, total tau (t-tau) and phosphorylated tau (p-tau). CSF concentrations levels were available from analyses performed as part of diagnostic work-up using clinically validated commercial immunoassays. Pathological status (normal vs. abnormal) followed laboratory reference cut-offs. Structural magnetic resonance imaging (MRI) scans of AD participants were evaluated using standardized visual rating scales to quantify regional atrophy and white matter hyperintensity (WMH) burden. Atrophy was rated according to the validated criteria described by Harper et al. (77) and the ARWMC scale, and performed by trained raters blinded to clinical information.
For the AD cohort, cognitive testing followed the standard CERAD-NAB+ protocol (78) and was complemented by additional measures: digit span (forward and backward), logical memory, and visual reproduction from the WMS-IV (79); block design from the WAIS-IV (80); and the alertness subtest of the Test of Attentional Performance (TAP) (81). Functional and neuropsychiatric symptoms were assessed using the Bayer ADL scale (82), the Neuropsychiatric Inventory (83), and standardized mood measures (Beck Depression Inventory, Geriatric Depression Scale, or HADS) (84–86). Normative adjustments followed each test’s requirements: CERAD-NAB+ was corrected for age, education, and sex; WMS-IV and WAIS-IV for age; and TAP for age and education. Cognitive impairment was defined as performance ≥1.5 SD below the respective normative values.
We recruited an equal-sized cognitively healthy control (CHC) group (N = 22) from the same geographical region as the AD participants. All controls provided written informed consent and had no history of major neurological or psychiatric disorders, including stroke, epilepsy, traumatic brain injury with loss of consciousness, neurodegenerative disease, psychotic or bipolar disorder, or substance use disorder. Inclusion required native German proficiency and the absence of persistent voice or speech disorders or untreated hearing problems. Controls were matched to the AD group on age, sex, and education. Characteristics for both groups are summarized in Table 1.
Table 1
| Variables | AD (N = 22) | CHC (N = 22) |
|---|---|---|
| Age at assessment (years) | 69.32 ± 7.29 | 70.12 ± 6.89 |
| Sex (female/male) | 11/11 | 11/11 |
| Education (ISCED level) | 3.50 ± 1.68 | 3.45 ± 1.57 |
| Clinical severity | ||
| Subjective cognitive decline (SCD) | 4 (18.2%) | – |
| Mild cognitive impairment (MCI) | 15 (68.2%) | – |
| Mild dementia | 3 (13.6%) | – |
| Cerebrospinal fluid biomarkers | – | |
| Aβ₁–₄₂ pathological status | 15 (68.2%) | – |
| Aβ₁–₄₂/Aβ₁–₄₀ ratio pathological status | 20 (90.9%) | – |
| Total tau (t-tau) pathological status | 16 (72.7%) | – |
| Phosphorylated tau (p-tau) pathological status | 19 (86.4%) | – |
| Clinical MRI visual ratings | – | |
| Orbitofrontal atrophy score | 1 (0) | – |
| Rostral anterior atrophy score | 2 (1) | – |
| Anterior temporal atrophy score | 1 (0.5) | – |
| Fronto-insular atrophy score | 1 (1) | – |
| Medial temporal atrophy score | 1 (1) | – |
| Posterior atrophy score | 2 (1) | – |
| ARWMC Basal ganglia score | 0 (0) | – |
| ARWMC Periventricular score | 1.5 (1) | – |
Demographic and clinical characteristics.
Values for continuous variables are presented as mean ± SD; MRI visual rating scale scores as median (IQR); clinical severity and biomarker categories as N (%). AD, Alzheimer’s disease; CHC, cognitively healthy controls; ARWMC, Age-Related White Matter Changes scale; ISCED, International Standard Classification of Education (159); MRI, magnetic resonance imaging; SCD, subjective cognitive decline; MCI, mild cognitive impairment.
2.2 Speech data elicitation
Speech samples were elicited using the Cookie Theft picture-description task from the Boston Diagnostic Aphasia Examination (BDAE) (87). The task requires participants to produce an unconstrained verbal narrative describing a complex household scene, providing a standardized method for sampling spontaneous, connected speech. All participants were instructed to describe the picture in as much detail as possible (“Tell me everything you see happening in this picture”), without time limits or additional prompting.
2.3 Linguistic biomarkers computation
Linguistic biomarkers were derived through a rigorous two-stage processing pipeline that first generated high-fidelity speech transcripts and then performed downstream automatic extraction of expert-engineered and clinically meaningful features. Speech recordings were transcribed using Whisper Large-v3, a state-of-the-art open-source automatic speech recognition (ASR) model developed by OpenAI (88). Whisper is a transformer-based architecture trained on a large, weakly supervised multilingual corpus and is characterized by strong robustness to background noise, speaker variability, and distributional shifts. We employed the German language configuration of the Large-v3 model. All automatically generated transcripts were subsequently reviewed by the authors. The transcribed speech samples served as input to CYMO, a next-generation text mining and analytics platform developed by Exaia Technologies.1 CYMO utilizes spaCy, an industrial-standard NLP library2, for fundamental NLP tasks, benefiting from its robustness, strong performance across established benchmarks, speed, and ease of deployment and maintenance. These NLP tasks include: (1) Tokenization, (2) Part-of-Speech (POS) Tagging, (3), Lemmatization (rule-based), (4) Dependency Parsing, and (5) Sentence Segmentation. As an integrated end-to-end solution, CYMO eliminates the limitations of fragmented linguistic-processing workflows by providing a seamless and unified environment for feature extraction. CYMO’s measurement module applies a sliding-window framework to compute context-sensitive distributions of linguistic properties across text segments, capturing fine-grained variation in phrasal, clausal, and sentential complexity. This architecture enables transparent, reproducible, and scalable extraction of interpretable linguistic biomarkers. The tutorial on CYMO with its comprehensive description is publicly available in a GitHub repository.3
The feature framework is grounded in multidisciplinary cognitive and behavioral neuroscience, which elucidates how humans acquire, develop, and process language and how these processes depend on working memory, executive function, and attentional control (89–91). Building on this multidisciplinary foundation, we next detail the three main categories of linguistic biomarkers incorporated in our framework—(1) syntactic complexity, (2) lexical density, diversity, and sophistication, and (3) information-theoretic metrics—and outline the specific theoretical constructs that motivate each.
Syntactic complexity—the diversity and sophistication of sentence structures—indexes core language acquisition and processing mechanisms across the lifespan (92–94). During development, children acquire progressively more complex constructions, and literacy further broadens the available syntactic repertoire. Comprehending and producing complex syntax impose demands on both language-specific and domain-general cognitive resources, particularly working memory and executive control. Consequently, syntactic complexity serves as a sensitive indicator of the integrity of underlying neurocognitive systems. This is especially pertinent in Alzheimer’s disease, where deficits in memory and executive control limit the generation, maintenance, and integration of complex structures. Individuals with Alzheimer’s characteristically produce shorter syntactic units (phrases/clauses/sentences), reduced subordination and coordination, and fewer complex phrasal constructions (44, 56, 95). A recent systematic review corroborates this pattern, reporting moderate to large associations between syntactic simplification and cognitive impairment (96). In our study, syntactic complexity was operationalized using four metrics that capture complementary aspects of structural organization: coordination, indexed by coordinate phrases per clause (cPC); length of production units, measured as mean clause length (MCL) and mean sentence length (MSL); and sentence complexity, quantified as clauses per sentence (CS).
Lexical complexity—the diversity, informativeness, and sophistication of the words produced in spontaneous language—indexes fundamental aspects of vocabulary knowledge and lexical access across the lifespan. It comprises three principal dimensions: lexical density, lexical diversity, and lexical sophistication. Lexical density is defined as the proportion of lexical items (e.g., nouns, verbs, adjectives) relative to the total number of words (tokens) in an utterance (97). Lexical diversity, also referred to as lexical variation, reflects the range of non-repetitive word forms in a speech sample, typically expressed as the number of unique types relative to the total token count (98). More precisely, it captures the production of phonologically and orthographically distinct word forms and serves as an index of the breadth of a speaker’s accessible vocabulary; common operationalizations quantify the type–token relationship using measures such as the type–token ratio (TTR) and length-adjusted variants including root TTR (rTTR) (99) and bilogarithmic TTR (bTTR) (100). Lexical sophistication captures the proportion of relatively infrequent, advanced, or more precise lexical items in a speech sample (101), typically quantified by comparing produced words to corpus-based frequency lists and calculating the share that falls within lower-frequency ranges. In the present work, lexical sophistication was quantified using corpus-based frequency measures that compute the mean log-frequency of all words in a speech sample—normalized by utterance length—based on their frequencies in academic, fiction, news, and spoken language corpora, with lower values indicating rarer and more advanced vocabulary. As a complementary indicator of lexical complexity, we also included mean length of word in characters (MLWc), which captures the use of morphologically and orthographically complex word forms.
Across the lifespan, lexical complexity metrics—particularly word frequency, familiarity, and lexical sophistication—exert robust effects on real-time processing: self-paced reading and eye-tracking studies consistently show that lower-frequency, less familiar, and more semantically specific words elicit longer reading times, increased fixation durations, and reduced skipping (102–104). These same lexical dimensions also differentiate clinical populations: individuals with Alzheimer’s disease typically exhibit reduced lexical richness, lower lexical density, and increased redundancy, often accompanied by “empty” or semantically underspecified expressions, compared with healthy controls (105, 106) [see also (54)]. For a recent synthesis of studies highlighting the moderate to high importance of these lexical metrics, see Shankar et al. (46).
Information-theoretic metrics—the third major category—capture the compressibility, predictability, and statistical organization of language and index core cognitive mechanisms involved in statistical learning, predictive processing, and efficient information encoding. These processes support how humans acquire, store, and deploy the probabilistic structure of language across the lifespan (89, 90), yet remain largely underexamined in Alzheimer’s disease research. The first subcategory, Kolmogorov complexity, quantifies the algorithmic structure of speech using compression-based approximations (here, Kolmogorov Deflate, a DEFLATE-derived measure combining LZ77 and Huffman coding), with lower complexity reflecting increased redundancy and reduced structural diversity (107–109)—patterns consistent with discourse simplification in Alzheimer’s disease (46, 110, 111). The second subcategory, predictive sequencing, is rooted in statistical learning and the brain’s predictive architecture: through lifelong exposure, speakers accumulate sensitivity to the frequency and distribution of multiword sequences (112–116). Robust multiword frequency effects are well documented in children acquiring their first language (117, 118), adult native speakers in self-paced reading and eye-tracking (119), and second-language learners, who show parallel facilitatory effects and individual differences modulated by working memory and other cognitive resources (120–122). In the present work, predictive sequencing was operationalized using Normalized Log Frequency (NLF) scores for bigrams, trigrams, fourgrams, and fivegrams across academic, fiction, news, and spoken registers, capturing the extent to which speakers rely on distributional patterns entrenched through experience. Although the connection between distributional learning and cognitive reserve is still theoretical, longitudinal studies such as the Nun Study highlight that richer and more complex lifelong language experiences contribute to reserve more broadly (57). Information-theoretic metrics offer a sensitive means of capturing impairments in the statistical and predictive foundations of language that remain undetected by syntactic or lexical features.
Across all three domains—syntactic complexity, lexical complexity, and information-theoretic metrics—32 linguistic biomarkers were extracted through CYMO, introduced earlier in this section. Using its implemented sliding-window technique, CYMO enables the computation of multiple measurements for each biomarker for every sentence or utterance, rather than over the entire speech sample. This produces robust, high-resolution estimates of language use. In contrast, conventional approaches typically generate only a single aggregate score per sample, limiting granularity and sensitivity. The extracted biomarkers are detailed in Table 2, which lists each feature’s name, code, category, subcategory, and a concise description of its computation.
Table 2
| # | Biomarker name | Code | Category | Subcategory | Description |
|---|---|---|---|---|---|
| 1 | Mean length of sentence | MLS | Syntactic Complexity | Length of production unit | Mean number of words per sentence. |
| 2 | Mean length of clause | MLC | Syntactic Complexity | Length of production unit | Mean number of words per clause. |
| 3 | Sentence complexity ratio | CS | Syntactic Complexity | Sentence complexity | Mean number of clauses per sentence. |
| 4 | Coordinate phrases per clause | cPC | Syntactic Complexity | Coordination | Mean number of coordinated phrases per clause. |
| 5 | Lexical density | LD | Lexical Richness | Lexical density | Proportion of content words to total words. |
| 6 | Number of different words | NDW | Lexical Richness | Lexical diversity | Total number of unique words. |
| 7 | Type-token ratio | TTR | Lexical Richness | Lexical diversity | Ratio of unique words to total words. |
| 8 | Bilogarithmic TTR | bTTR | Lexical Richness | Lexical diversity | Ratio of log unique words to log tokens. |
| 9 | Root TTR | rTTR | Lexical Richness | Lexical diversity | Unique words divided by √tokens. |
| 10 | Corrected TTR | cTTR | Lexical Richness | Lexical diversity | Unique words divided by √(2 × tokens). |
| 11 | Mean length of word (characters) | MLWc | Lexical Richness | Lexical sophistication | Mean number of characters per word. |
| 12–15 | Unigram NLF (academic, fiction, news, spoken) | 1GNLFa–1GNLFs | Lexical Richness | Lexical sophistication | Normalized Log Frequency of unigrams vs. respective domain corpus. |
| 16 | Kolmogorov Deflate | KDbase | Information-Theoretic | Kolmogorov complexity | Compression ratio (compressed/original). |
| 17–32 | Bigram–Fivegram NLF (academic, fiction, news, spoken) | 2GNLFa–5GNLFs | Information-Theoretic | Predictive sequencing | Normalized Log Frequency of n-grams vs. respective domain corpus. |
Linguistic biomarkers and their classification.
* NLF (Normalized Log Frequency) quantifies alignment of speech n-grams with reference corpus frequency distributions. Descriptions for all NLF measures follow the pattern: “NLF of n-grams vs. corpus”.
2.4 Machine-learning framework for AD classification
We implemented supervised binary classification models for automatic detection of CSF-validated AD from connected speech. The models distinguish biomarker-positive AD cases from cognitively healthy, demographically matched controls and are trained on theory-driven, clinically meaningful linguistic biomarkers described in detail in Section 2.3. The framework was designed to maximize statistical robustness, guard against overfitting, and enable transparent interpretation. The pipeline comprised four sequential components: (1) stability-based feature selection to identify a subset of linguistic biomarkers that exhibit consistent discriminative value across resampled training folds; (2) model training and validation based on stratified cross-validation with strict separation of training and test data; (3) performance evaluation using complementary metrics that jointly characterize classification performance; and (4) post hoc model interpretation using a model-agnostic feature-attribution approach to quantify the magnitude and direction of individual feature contributions to AD–control discrimination.
2.4.1 Stability-based feature selection
Feature selection was applied to the full set of 32 CYMO-derived linguistic biomarkers using a two-stage, stability-based recursive feature elimination (RFE) framework, consistent with recent approaches in machine-learning and clinical prediction settings (123, 124). In the first stage, RFE with a logistic regression estimator was applied separately within each training set of the stratified five-fold cross-validation, iteratively discarding the least informative features as indexed by the absolute magnitude of model coefficients. In the second stage, selection frequencies were aggregated across folds, and features retained in at least three of five folds (≥60%) were defined as stable predictors.
2.4.2 Model training and validation
Five supervised classification algorithms were implemented: Logistic Regression (LR), Support Vector Machine with a radial-basis-function kernel (SVM-RBF), Random Forest (RF), Gradient Boosting Decision Trees (GBDT), and Extreme Gradient Boosting (XGBoost). LR, SVM-RBF, RF, and GBDT were implemented using scikit-learn (125), and XGBoost using the XGBoost Python library (126). These algorithms are widely used in clinical prediction tasks, including Alzheimer’s disease, owing to their robustness and interpretability (46, 61, 96). Together, these models span regularized linear classifiers, kernel-based methods, and tree-based ensembles, enabling a direct comparison of classification performance across distinct model families.
To evaluate model performance and prevent overfitting, we trained all classifiers using five-fold stratified cross-validation. Each fold preserved the class distribution between Alzheimer’s disease (AD) and cognitively healthy control (CHC) participants. For each split, imputation and feature scaling were fit on the training data and applied to the held-out fold to avoid data leakage. All classifiers were trained on the stability-selected linguistic feature set defined in Section 2.4.1. Hyperparameters were set according to the configurations specified in Table 3 and held constant across all cross-validation folds.
Table 3
| Model | Key settings |
|---|---|
| Logistic regression | max_iter = 3,000, class_weight = “balanced,” C = 0.5, solver = library default |
| SVM (RBF) | kernel = “rbf,” probability = True, class_weight = “balanced,” C = 1.0, gamma = “scale” |
| Random forest | n_estimators = 200, max_depth = 4, class_weight = “balanced,” other params = defaults |
| Gradient boosting (GBDT) | n_estimators = 200, learning_rate = 0.05, max_depth = 3, other params = defaults |
| XGBoost | n_estimators = 200, learning_rate = 0.05, max_depth = 3, eval_metric = “logloss,” other params = defaults |
Hyperparameter settings for all classifiers.
2.4.3 Evaluation metrics
Model performance was evaluated using standard metrics for binary classification, including accuracy, sensitivity (recall), specificity, precision, F1 score, Matthews correlation coefficient (MCC), the area under the receiver operating characteristic curve (AUROC), and the area under the precision–recall curve (AUPRC) (Equations 1–5). AD diagnosis was treated as the positive class, and the F1 score for AD was used as the primary performance metric, with the remaining measures providing complementary summaries of discrimination between AD patients and CHC. Confusion matrices were aggregated across cross-validation folds to visualize classification consistency and error patterns. To quantify uncertainty in this small-sample setting, 95% bootstrap confidence intervals were computed by resampling cross-validation folds with replacement (1,000 iterations) and recomputing the mean metric value at each iteration. For completeness, the metrics are defined as follows, with TP, TN, FP, and FN denoting true positives, true negatives, false positives, and false negatives, respectively (AD as the positive class).
Accuracy measures the overall proportion of correctly classified observations across both classes.
The F1-score is the harmonic mean of precision and recall, emphasizing balanced control of false positives and false negatives.
where:
Sensitivity (Recall, True Positive Rate) quantifies the proportion of AD cases correctly identified by the model. High sensitivity indicates effective detection of true disease cases, minimizing missed diagnoses.
Specificity (True Negative Rate) measures the proportion of CHC participants correctly recognized as non-AD. High specificity reflects a low false-positive rate and reliable exclusion of cognitively normal individuals.
Matthews Correlation Coefficient (MCC) represents the correlation between predicted and true labels, ranging from −1 (total disagreement) to +1 (perfect prediction). It is regarded as one of the most balanced single-number measures of binary classifier quality.
2.4.3.1 Area under the receiver operating characteristic curve
AUROC quantifies the probability that a randomly chosen AD case receives a higher predicted probability than a randomly chosen CHC. It summarizes the trade-off between sensitivity and 1 – specificity across all decision thresholds. Values range from 0.5 (chance-level discrimination) to 1.0 (perfect separation).
2.4.3.2 Area under the precision–recall curve
AUPRC provides a threshold-independent summary of how well the model maintains precision as recall increases and complements AUROC by explicitly characterizing the precision–recall trade-off for the AD class across all decision thresholds.
2.4.4 Model interpretability
Model interpretability is a key prerequisite for responsible AI and a necessary condition for the safe deployment of prediction models in clinical practice (58). This requirement extends to speech- and language-based biomarker models for Alzheimer’s disease (AD), where machine-learning systems are used to distinguish AD patients from CHC. In this study, interpretability is defined as the ability to attribute model predictions to specific linguistic input features in a stable and clinically meaningful manner, both at the level of individual cases (local explanations) and when aggregated across the cohort (global feature importance). Feature importance was quantified using Shapley values, computed in Python with the shap library in accordance with the SHAP (SHapley Additive exPlanations) framework (127–129), which decomposes each prediction into an additive set of feature-wise contributions. Global interpretability was obtained by summarizing SHAP values across participants to characterize which linguistic biomarkers are most strongly associated with AD versus cognitively healthy aging, whereas local interpretability was achieved by examining participant-level SHAP profiles, with particular emphasis on AD patients, to determine which features drive individual predictions and to identify potentially implausible model behavior. This SHAP-based attribution supports responsible AI by linking model behavior to clinically interpretable variables in a transparent and reproducible way (130).
3 Results
Using the machine learning framework described in Section 2.4, we first report the outcome of the stability-based feature selection applied to the full set of CYMO-derived linguistic metrics (Section 3.1), then evaluate cross-validated AD–CN classification performance obtained from the resulting predictor set (Section 3.2), and examine SHAP-based feature attributions at global (cohort-level) and local (participant-level) levels (Section 3.3).
3.1 Stability-based feature selection outcomes
Applying the stability-based RFE framework (Section 2.4.1) to the 32 CYMO-derived linguistic markers yielded a parsimonious set of seven stable predictors. Aggregating selection frequencies across the five stratified folds and applying the ≥60% stability threshold (retained in at least three of five folds) reduced the feature space to: mean length of sentence (MLS), mean length of clause (MLC), lexical density (LD), Kolmogorov Deflate (KDbase), and three 4-gram Normalized Log Frequency measures in the academic, news, and spoken domains (4GNLFa, 4GNLFn, 4GNLFs).
In accordance with the predefined categories in Table 2, the stability-selected predictors were drawn from all three linguistic domains. The information-theoretic domain contributed four predictors, spanning both of its subcategories—Kolmogorov complexity (KDbase) and predictive sequencing (4GNLFa, 4GNLFn, 4GNLFs). The syntactic complexity domain contributed two predictors from the length-of-production-unit subcategory (MLS, MLC), and the lexical richness domain contributed one predictor from the lexical-density subcategory (LD). All other markers fell below the stability threshold and were not included in subsequent modeling or interpretability analyses.
3.2 Performance of AD classification models
Table 1 summarizes stratified 5-fold cross-validation performance for all five classifiers trained on the stability-selected linguistic predictors. Across models, F1-scores ranged from 0.881 to 0.900 and accuracies from 0.889 to 0.911. AUROC values ranged from 0.905 to 0.990 and AUPRC from 0.878 to 0.990, indicating strong discrimination between AD and CHC participants. All metrics are reported as mean ± standard deviation across cross-validation folds.
Logistic regression (LR) achieved the highest performance (F1 = 0.900 ± 0.133; accuracy = 0.911 ± 0.109; MCC = 0.839 ± 0.197), together with the largest area measures (AUROC = 0.990 ± 0.020; AUPRC = 0.990 ± 0.020) and a symmetric error profile (sensitivity = 0.900 ± 0.200; specificity = 0.900 ± 0.200). SVM-RBF showed similar F1 (0.900 ± 0.133) and accuracy (0.911 ± 0.109), with slightly lower AUROC (0.980 ± 0.040) and AUPRC (0.978 ± 0.045). Tree-based methods (random forest, gradient boosting, XGBoost) reached comparable accuracies and F1-scores (F1 = 0.881–0.896; accuracy = 0.889–0.911) but exhibited lower MCC values (0.783–0.821) and area measures (AUROC = 0.930–0.966; AUPRC = 0.878–0.966), as well as larger standard deviations across folds (Table 1).
ROC and precision–recall curves derived from the pooled out-of-fold predictions are shown in Figure 1. The curve profiles closely reflect the performance differences reported in Table 1. Logistic regression shows the strongest separation between AD and CHC, with consistently higher ROC and PR curves across the full threshold range. SVM-RBF shows similarly high performance, followed by random forest, which displays a modest reduction in precision at intermediate recall levels. Gradient boosting and XGBoost yield lower but still clearly discriminative curves in both panels. Consistent with the AUROC and AUPRC values, logistic regression achieves the largest areas and the most stable performance across folds and was therefore selected as the reference model for subsequent SHAP-based feature attribution analyses (Table 4).
Figure 1
Table 4
| Model | Accuracy | F1 | Sensitivity | Specificity | MCC | AUROC | AUPRC |
|---|---|---|---|---|---|---|---|
| Logistic regression | 0.911 ± 0.109 | 0.900 ± 0.133 | 0.900 ± 0.200 | 0.900 ± 0.200 | 0.839 ± 0.197 | 0.990 ± 0.020 | 0.990 ± 0.020 |
| SVM (RBF) | 0.911 ± 0.109 | 0.900 ± 0.133 | 0.900 ± 0.200 | 0.900 ± 0.200 | 0.839 ± 0.197 | 0.980 ± 0.040 | 0.978 ± 0.045 |
| Random forest | 0.911 ± 0.130 | 0.896 ± 0.166 | 0.900 ± 0.200 | 0.910 ± 0.111 | 0.821 ± 0.265 | 0.960 ± 0.080 | 0.966 ± 0.068 |
| Gradient boosting | 0.911 ± 0.130 | 0.896 ± 0.166 | 0.900 ± 0.200 | 0.910 ± 0.111 | 0.821 ± 0.265 | 0.905 ± 0.136 | 0.878 ± 0.174 |
| XGBoost | 0.889 ± 0.141 | 0.881 ± 0.168 | 0.900 ± 0.200 | 0.860 ± 0.196 | 0.783 ± 0.281 | 0.930 ± 0.140 | 0.911 ± 0.178 |
Cross-validated performance of five classifiers trained to distinguish Alzheimer’s disease (AD) from cognitively healthy controls (CHC).
Values are mean ± SD across stratified 5-fold cross-validation.
The confusion matrix for the logistic regression model is shown in Figure 2. It contains four boundary misclassifications—two AD and two CHC—with predicted AD probabilities in the intermediate range (0.38–0.65), indicating proximity to the decision boundary. The symmetric distribution of errors reflects balanced classifier behavior.
Figure 2
3.3 Model interpretability
The SHAP summary plots in Figure 3 show how the seven stability-selected linguistic predictors contribute to the logistic-regression model’s prediction of Alzheimer’s disease. SHAP values were computed at the local level for each participant, quantifying each predictor’s contribution to the model output relative to the baseline expectation. Panel A displays the distribution of SHAP values across predictors, and Panel B reports global importance based on mean absolute SHAP values.
Figure 3
Rankings based on mean absolute SHAP values showed that KDbase had the largest contribution (~31%), followed by lexical density (~14%), mean clause length (~13%), and mean sentence length (~12%). The three four-gram normalized log-frequency measures accounted for the remaining ~28% of total attribution (4GNLFs ~ 10%, 4GNLFa ~ 9%, 4GNLFn ~ 9%). These predictors reflect all three linguistic categories defined in Section 2.3: information-theoretic metrics (Kolmogorov complexity; predictive sequencing), syntactic complexity (length of production units), and lexical richness (lexical density). Higher mean absolute SHAP values indicate a greater influence on the model’s output.
Across all predictors, lower values were consistently associated with positive SHAP contributions, corresponding to a higher model-predicted probability of Alzheimer’s disease. This pattern indicates that participants with Alzheimer’s disease produced more compressible and less information-dense speech (lower KDbase), reduced lexical content (lower lexical density), shorter syntactic units (lower mean clause and sentence length), and fewer statistically entrenched multiword combinations (lower four-gram normalized log-frequency scores). Together, these characteristics define the predictor profile most strongly shifting model output toward the AD class.
A local explanation for an individual participant correctly classified as having Alzheimer’s disease is shown in Figure 4. The waterfall plot decomposes the model output into additive contributions from the seven linguistic predictors, starting at the baseline expectation (E[f(x)] = −0.13) and summing to the final log-odds (f(x) = 3.88). The largest positive shifts toward an AD prediction were produced by low KDbase (+1.79) and low lexical density (+1.20). Additional positive contributions arose from short clause length (+0.48) and from lower four-gram normalized log-frequency scores in the academic (+0.33), spoken (+0.09), and news (+0.06) domains, as well as short sentence length (+0.06). No substantial negative contributions were observed, yielding a prediction strongly favoring the AD class. This local explanation reflects the same pattern seen in the cohort-level attributions: lower information-theoretic, lexical, and syntactic values jointly shift the model output toward an Alzheimer’s disease classification.
Figure 4
4 Discussion
We have demonstrated the feasibility of predicting Alzheimer’s disease from spontaneous speech in the German language using an end-to-end, linguistically interpretable machine-learning framework. Across five supervised classifiers trained on a compact set of seven stability-selected linguistic biomarkers, cross-validated performance was consistently high (mean accuracy ≈0.91; mean F1 ≈ 0.90). To our knowledge, this is the first study to detect Alzheimer’s disease from spontaneous German speech in a cohort whose diagnostic status is defined according to clinical-biological criteria anchored in CSF biomarkers and directly compared with cognitively normal controls carefully matched on age, sex, and education. These findings indicate that connected speech in German carries a strong and compact linguistic signature of biomarker-confirmed AD that can be captured reliably by an end-to-end model. By combining a standardized Cookie Theft picture-description task with a demographically balanced design, our study follows the methodological principles exemplified by ADReSS and ADReSSo (62, 63), which have shown how standardized elicitation and carefully matched Alzheimer’s and control groups, together with a shared high-quality dataset, can act as a catalyst for progress in speech-based AD detection by enabling diverse machine-learning approaches to be developed, tested, and directly compared on a common benchmark. Our work extends these best-practice principles beyond English, contributing to the broader move towards cross-linguistic, standardized datasets (76). We further encourage future studies to construct demographically and socioeconomically balanced datasets in additional languages and settings to minimize non-disease variance. In parallel, our focus on connected spontaneous speech rather than isolated word lists or single-sentence recall is consistent with evidence that continuous, naturalistic speech production provides a more sensitive and ecologically valid marker of Alzheimer-related decline than traditional neuropsychological tasks (55, 110).
Beyond predictive performance, a key objective was to align our end-to-end ML framework with emerging desiderata for responsible AI-driven speech biomarkers required for clinical adoption and decreasing timelines to translation (58). In this context, our results demonstrate that robust discrimination between biomarker-confirmed AD and matched controls can be achieved without resorting to opaque, high-dimensional latent representations that dominate much of the current literature (61, 131–134). By constraining modeling to a low-dimensional, stability-selected set of seven predictors, we show that high accuracy is attainable in a small clinical sample while reducing the risk that performance is driven by idiosyncrasies of the training cohort rather than disease-related signal. At the same time, the framework directly addresses the second major limitation of black-box approaches, namely their lack of alignment with the clinical presentation and characteristic linguistic deficits of AD, by operating on expert-engineered, theoretically grounded, and individually validated linguistic markers and combining these with SHAP-based attribution to yield subject-level explanations that link each prediction to specific properties of a patient’s speech.
Taken together, the resulting stability-selected feature set provides a concise yet mechanistically informative characterization of AD-related changes in spontaneous speech: shorter clauses and sentences, reduced lexical density, and lower information-theoretic complexity and predictive sequencing. This pattern is consistent with English-language evidence of syntactic simplification, impoverished vocabulary, and reduced informational content in the speech of patients with AD (46, 60, 76, 110, 132), while the inclusion of Kolmogorov-based and n-gram–based metrics indicates that disturbances in information-theoretic complexity and predictive sequencing constitute an additional, disease-relevant dimension not fully captured by traditional syntactic or lexical measures. This is consonant with work showing that information-theoretic measures derived from probabilistic language models systematically covary with activity in fronto–temporal networks during naturalistic language comprehension (113) and with large-scale neural dynamics underlying internally generated language and thought (135).
From a systems-neuroscience perspective, the speech-based digital AD phenotype delineated here is congruent with the known organization and vulnerability of large-scale neural networks in AD. Connected-speech production engages a distributed fronto–temporal–parietal language network together with medial temporal memory and frontoparietal control systems that jointly support lexical–semantic retrieval, syntactic planning, and maintenance of message-level content over time [for a recent meta-analytic connectivity modeling study, see Hsu et al. (136)]. Neuroimaging and network-analytic studies of AD demonstrate early and progressive disruption of these default-mode, temporal–parietal language, medial temporal, and control networks (137–140). Within this framework, the observed pattern of shortened clauses and sentences, reduced lexical density, and altered information-theoretic organization can be interpreted as the behavioral manifestation of a diminished capacity of these circuits to support complex, information-rich connected speech.
Our work contributes to a limited but emerging literature linking speech-based digital markers to CSF measures of amyloid and tau pathology. Across English, Spanish, and Mandarin cohorts, prior studies have shown that lexical features, acoustic parameters, and aggregate speech indices are associated with CSF concentrations of amyloid and tau, both in early Alzheimer’s disease and in cognitively unimpaired at-risk populations (141–146). However, this literature has largely focused on amyloid status or continuous CSF biomarker levels in heterogeneous preclinical or prodromal samples, has relied primarily on acoustic or relatively coarse speech descriptors, and has rarely produced theory-driven machine-learning approaches calibrated to CSF-anchored, clinically diagnosed Alzheimer’s disease.
Within this CSF-anchored literature, the present study extends existing work in three main respects. First, we show that spontaneous connected speech in German supports high-accuracy discrimination between Alzheimer’s disease and matched controls when diagnostic status is defined in accordance with contemporary International Working Group recommendations, using a CSF panel comprising Aβ₁–₄₂ concentration, the Aβ₁–₄₂/Aβ₁–₄₀ ratio, total tau (t-tau), and phosphorylated tau (p-tau). Second, rather than modeling amyloid positivity or raw CSF biomarker concentrations in heterogeneous at-risk or preclinical cohorts, we target clinically manifest, CSF-verified Alzheimer’s disease and directly contrast these patients with cognitively normal controls rigorously matched on age, sex, and education, building on the matching strategy outlined above to minimize non-disease variance. Third, we couple this strict biological reference standard with a compact, theory-driven, linguistically interpretable feature set and SHAP-based attribution within an end-to-end framework, yielding an empirically validated speech-based Alzheimer phenotype and subject-level explanations that link individual predictions to specific properties of connected speech, as detailed in the preceding sections.
5 Limitations
Several limitations of this work should be acknowledged when interpreting and generalizing our findings. First, the sample size is modest (22 CSF-verified AD cases and 22 matched controls). While this is typical for deeply phenotyped, biomarker-anchored cohorts and sufficient to support the internal cross-validated analyses reported here, it constrains inferences about robustness across the broader clinical population. Future studies will need to replicate and extend these results in larger samples that span a wider range of ages, educational backgrounds, dialects, comorbidities, and clinical stages, to establish the stability of the identified speech phenotype under more heterogeneous real-world conditions. Second, future work should develop fusion models that integrate the linguistically interpretable markers used here with paralinguistic fluency markers—particularly pause behavior, hesitations, speech rate, and articulation rate—which have been shown in English-language studies to enhance dementia classification when combined with text-based features (72, 147–149). In addition, the syntactic complexity metrics used in the present study represent relatively coarse structural indicators derived from surface syntactic properties; future work should incorporate more advanced dependency-based measures such as T-units and complex nominals (150–154). Third, our current models do not yet incorporate information about individual personality profiles, despite converging evidence that stable traits play an important role in AD risk and trajectories of cognitive and brain aging (155). For example, studies have shown that higher neuroticism is associated with increased dementia risk, whereas higher conscientiousness and openness to experience may confer relative protection from dementia incidence (156, 157). Building on recent information-fusion and multitask-learning approaches that combine language-derived signals with personality profiles (158), an important direction for future work will be to extend our framework to jointly model speech-based Alzheimer’s markers and personality traits, in order to evaluate whether trait-like vulnerability profiles add incremental value for detection and risk stratification. Fourth, the present analyses are restricted to cross-sectional assessments and thus quantify diagnostic discrimination at a single time point. Future analyses of longitudinal data could complement the characterization of individual trajectories of cognitive and functional deterioration, track within-patient disease progression, and assess sensitivity to treatment- or prevention-related change over time.
6 Conclusion
Spontaneous speech offers a uniquely scalable window into the cognitive and neural systems most vulnerable in AD, and the present work shows that this window can be determined with clinically meaningful precision. Using German connected speech and an interpretable end-to-end machine-learning framework in a group of patients with a clinical-biological diagnose of AD, we identify a compact, mechanistically informative constellation of linguistic markers that reliably discriminates biomarker-confirmed AD from demographically matched controls. The resulting speech-based digital phenotype is biologically grounded, in line with previous cross-linguistic evidence on Alzheimer-related language changes, and applicable to subject-level analyses through transparent feature engineering and model attribution. These properties directly address key prerequisites for responsible clinical translation, namely standardization, interpretability, and robustness of measures, and position AI-enabled speech markers as a promising, non-invasive, and cost-efficient complement to established diagnostic pathways for earlier detection, refined risk stratification, and longitudinal monitoring of AD in both research and clinical settings.
Statements
Data availability statement
The datasets presented in this article are not readily available because of data sharing restrictions imposed by the informed consent and current privacy and data protection legislations. Requests to access the datasets should be directed to acosta@ukaachen.de.
Ethics statement
The study was approved by the Ethics Committee of the Faculty of Medicine of the RWTH Aachen University (EK 384/20) and is registered in the German Clinical Trials Register (DRKS00030633). The studies were conducted in accordance with the local legislation and institutional requirements. The participants provided their written informed consent to participate in this study.
Author contributions
DW: Conceptualization, Formal analysis, Methodology, Visualization, Writing – original draft, Writing – review & editing. EK: Conceptualization, Formal analysis, Methodology, Supervision, Writing – original draft, Writing – review & editing. MA: Data curation, Investigation, Writing – review & editing. YQ: Data curation, Software, Validation, Writing – review & editing. JP: Data curation, Project administration, Validation, Writing – review & editing. KR: Resources, Writing – review & editing. AC: Data curation, Funding acquisition, Investigation, Project administration, Supervision, Writing – review & editing.
Funding
The author(s) declared that financial support was received for this work and/or its publication. This research project was funded by the START-Program (grant number 118/20) of the Faculty of Medicine RWTH Aachen University and supported by the Brain Imaging Facility of the Interdisciplinary Center for Clinical Research Aachen within the Faculty of Medicine at RWTH Aachen University. No commercial funding was received for this work.
Conflict of interest
EK and YQ were employed by company Exaia Technologies. KR and AC were employed by Juelich Research Center GmbH.
The remaining author(s) declare that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The author(s) declared that Generative AI was not used in the creation of this manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Footnotes
1.^Beta version currently available for German: https://exaia-tech.com/bookdemo.
References
1.
World Health Organization (2025) Dementia World Health Organization. Available online at: https://www.who.int/news-room/fact-sheets/detail/dementia (Accessed March 24, 2026).
2.
ScheltensPDe StrooperBKivipeltoMHolstegeHChételatGTeunissenCEet al. Alzheimer's disease. Lancet. (2021) 397:33667416:1577–90. doi: 10.1016/S0140-6736(20)32205-4
3.
Alzheimer’s Disease International (2024) Dementia Statistics. Available online at: https://www.alzint.org/about/dementia-facts-figures/dementia-statistics/ (Accessed March 24, 2026).
4.
BatemanRJXiongCBenzingerTLSFaganAMGoateAFoxNCet al. Clinical and biomarker changes in dominantly inherited Alzheimer's disease. N Engl J Med. (2012) 367:795–804. doi: 10.1056/NEJMoa1202753,
5.
VillemagneVLBurnhamSBourgeatPBrownBEllisKASalvadoOet al. Amyloid β deposition, neurodegeneration, and cognitive decline in sporadic Alzheimer's disease: a prospective cohort study. Lancet Neurol. (2013) 12:357–67. doi: 10.1016/S1474-4422(13)70044-9,
6.
SelkoeDJHardyJ. The amyloid hypothesis of Alzheimer’s disease at 25 years. EMBO Mol Med. (2016) 8:595–608. doi: 10.15252/emmm.201606210
7.
van DyckCHSwansonCJAisenPBatemanRJChenCGeeMet al. Lecanemab in early Alzheimer’s disease. N Engl J Med. (2023) 388:9–21. doi: 10.1056/NEJMoa2212948
8.
SimsJRZimmerJAEvansCDLuMArdayfioPSparksJet al. Donanemab in early symptomatic Alzheimer disease. N Engl J Med. (2023) 388:1691–704. doi: 10.1001/jama.2023.13239
9.
DuboisBFeldmanHHJacovaCHampelHMolinuevoJLBlennowKet al. Advancing research diagnostic criteria for Alzheimer's disease: the IWG-2 criteria. Lancet Neurol. (2014) 13:614–29. doi: 10.1016/s1474-4422(14)70090-0
10.
DuboisBVillainNFrisoniGBRabinoviciGDSabbaghMCappaSet al. Clinical diagnosis of Alzheimer's disease: recommendations of the international working group. Lancet Neurol. (2021) 20:484–96. doi: 10.1016/S1474-4422(21)00066-1,
11.
DuboisBVillainNSchneiderLFoxNCampbellNGalaskoDet al. Alzheimer disease as a clinical–biological construct: an international working group recommendation. JAMA Neurol. (2024) 81:1304–11. doi: 10.1001/jamaneurol.2024.3770,
12.
HampelHVergalloAPerryGListaS. The Alzheimer precision medicine initiative. J Alzheimer's Dis. (2019) 68:1–24. doi: 10.3233/JAD-181121,
13.
MárquezFYassaMA. Neuroimaging biomarkers for Alzheimer's disease. Mol Neurodegener. (2019) 14:21. doi: 10.1186/s13024-019-0325-5,
14.
JackCRJrKnopmanDSJagustWJShawLMAisenPSWeinerMWet al. Hypothetical model of dynamic biomarkers of the Alzheimer's pathological cascade. Lancet Neurol. (2010) 9:119–28. doi: 10.1016/S1474-4422(09)70299-6,
15.
HardyJSelkoeDJ. The amyloid hypothesis of Alzheimer’s disease: Progress and problems on the road to therapeutics. Science. (2002) 297:353–6. doi: 10.1126/science.1072994,
16.
BraakHBraakE. Neuropathological stageing of Alzheimer-related changes. Acta Neuropathol. (1991) 82:239–59. doi: 10.1007/BF00308809,
17.
Spires-JonesTLHymanBT. The intersection of amyloid beta and tau at synapses in Alzheimer's disease. Neuron. (2014) 82:756–71. doi: 10.1016/j.neuron.2014.05.004,
18.
BraakHDel TrediciK. The preclinical phase of the pathological process underlying sporadic Alzheimer's disease. Brain. (2015) 138:2814–33. doi: 10.1093/brain/awv236,
19.
ZhangXYYangZLLuGMYangGFZhangLJ. PET/MR imaging: new frontier in Alzheimer’s disease and other dementias. Front Mol Neurosci. (2017) 10:343. doi: 10.3389/fnmol.2017.00343,
20.
BlennowKZetterbergH. Biomarkers for Alzheimer's disease: current status and prospects for the future. J Intern Med. (2018) 284:643–63. doi: 10.1111/joim.12816,
21.
OlssonBLautnerRAndreassonUÖhrfeltAPorteliusEBjerkeMet al. CSF and blood biomarkers for the diagnosis of Alzheimer's disease: a systematic review and meta-analysis. Lancet Neurol. (2016) 15:673–84. doi: 10.1016/s1474-4422(16)00070-3
22.
HanssonO. Biomarkers for neurodegenerative diseases. Nat Med. (2021) 27:954–63. doi: 10.1038/s41591-021-01382-x,
23.
JanelidzeSZetterbergHMattssonNPalmqvistSVandersticheleHLindbergOet al. CSF Aβ42/Aβ40 and Aβ42/Aβ38 ratios: better diagnostic markers of Alzheimer disease. Acta Neuropathol. (2016) 131:715–28. doi: 10.1002/acn3.274
24.
HanssonOLehmannSOttoMZetterbergHLewczukP. Advantages and disadvantages of the use of the CSF aβ₁–₄₂/aβ₁–₄₀ ratio in clinical and research settings. Alzheimer's Res Ther. (2019) 11:34. doi: 10.1186/s13195-019-0485-0
25.
LeuzyABollackAPellegrinoDTeunissenCELa JoieRRabinoviciGDet al. Considerations in the clinical use of amyloid PET and CSF biomarkers for Alzheimer's disease. Nat Rev Neurol. (2025). doi: 10.1002/alz.14528
26.
NakamuraAKanekoNVillemagneVLKatoTDoeckeJDoréVet al. High performance plasma amyloid-β biomarkers for Alzheimer's disease. Nature. (2018) 554:249–54. doi: 10.1038/nature25456,
27.
PalmqvistSJanelidzeSQuirozYTZetterbergHLoperaFStomrudEet al. Discriminative accuracy of plasma p-tau217 for Alzheimer disease vs other neurodegenerative disorders. JAMA. (2020) 324:772–81. doi: 10.1001/jama.2020.12134
28.
LeuzyAMattsson‐CarlgrenNPalmqvistSJanelidzeSDageJLHanssonOet al. Blood-based biomarkers for Alzheimer's disease. EMBO Mol Med. (2022) 14:e14781. doi: 10.15252/emmm.202114408
29.
SchindlerSEGalaskoDPereiraACRabinoviciGDSallowaySSuárez-CalvetMet al. Acceptable performance of blood biomarker tests of amyloid pathology—recommendations from the global CEO initiative on Alzheimer's disease. Nat Rev Neurol. (2024) 20:426–39. doi: 10.1038/s41582-024-00977-5,
30.
SperlingRADonohueMCRissmanRAJohnsonKARentzDMGrillJDet al. Amyloid and tau prediction of cognitive and functional decline in unimpaired older individuals: longitudinal data from the A4 and LEARN studies. J Prev Alzheimers Dis. (2024) 11:802–13. doi: 10.14283/jpad.2024.122,
31.
van der KallLMDoréVBourgeatPRoweCCVillemagneVLSalvadoOet al. Association of β-amyloid level, clinical progression, and longitudinal cognitive change in normal older individuals. Neurology. (2021) 96:e662–70. doi: 10.1212/WNL.0000000000011222,
32.
WelshKAButtersNMohsRCBeeklyDEdlandSFillenbaumG. The consortium to establish a registry for Alzheimer’s disease (CERAD). Part V. A normative study of the neuropsychological battery. Neurology. (1994) 44:609–14. doi: 10.1212/WNL.44.4.609
33.
NasreddineZSPhillipsNABédirianVCharbonneauSWhiteheadVCollinIet al. The Montreal cognitive assessment (MoCA): a brief screening tool for mild cognitive impairment. J Am Geriatr Soc. (2005) 53:695–9. doi: 10.1111/j.1532-5415.2005.53221.x,
34.
FolsteinMFFolsteinSEMcHughPR. “Mini-mental state”: a practical method for grading the cognitive state of patients for the clinician. J Psychiatr Res. (1975) 12:189–98. doi: 10.1016/0022-3956(75)90026-6,
35.
SpoonerDMPachanaNA. Ecological validity in neuropsychological assessment: a case for greater consideration in research with neurologically intact populations. Arch Clin Neuropsychol. (2006) 21:327–37. doi: 10.1016/j.acn.2006.04.004,
36.
KueperJKSpeechleyMMontero-OdassoM. The Alzheimer’s disease assessment scale–cognitive subscale (ADAS-cog): modifications and responsiveness in pre-dementia populations: a narrative review. J Alzheimer's Dis. (2018) 63:423–44. doi: 10.3233/JAD-170991,
37.
CalamiaMMarkonKTranelD. Scoring higher the second time around: Meta-analyses of practice effects in neuropsychological assessment. Clin Neuropsychol. (2012) 26:543–70. doi: 10.1080/13854046.2012.680913,
38.
DagumP. Digital biomarkers of cognitive function. NPJ Digit Med. (2018) 1:10. doi: 10.1038/s41746-018-0018-4,
39.
European Medicines Agency. Questions and Answers: Qualification of Digital Technology-Based Methodologies to Support Approval of Medicinal Products (EMA/219860/2020).European Medicines Agency (2020).
40.
KourtisLCRegeleOBWrightJMJonesGB. Digital biomarkers for Alzheimer's disease: the mobile/wearable devices opportunity. NPJ Digit Med. (2019) 2:9. doi: 10.1038/s41746-019-0084-2,
41.
InselTR. Digital phenotyping: technology for a new science of behavior. JAMA. (2017) 318:1215–6. doi: 10.1001/jama.2017.11295,
42.
PowellDPrasadABoschIReiterJPriceND. Walk, talk, think, see and feel: harnessing the power of digital biomarkers in healthcare. NPJ Digit Med. (2024) 7:4. doi: 10.1038/s41746-024-01023-w
43.
ZhangYWangJZongHSinglaRKUllahALiuXet al. The comprehensive clinical benefits of digital phenotyping: from broad adoption to full impact. NPJ Digit Med. (2025) 8:19. doi: 10.1038/s41746-025-01602-5,
44.
GumusMKooMStudzinskiCMBhanARobinJBlackSE. Linguistic changes in neurodegenerative diseases relate to clinical symptoms. Front Neurol. (2024) 15:1373341. doi: 10.3389/fneur.2024.1373341,
45.
CaoFVogelAPGharahkhaniPRenteriaME. Speech and language biomarkers for Parkinson's disease prediction, early diagnosis and progression. NPJ Parkinsons Dis. (2025) 11:57. doi: 10.1038/s41531-025-00913-4,
46.
ShankarRBundeleAMukhopadhyayA. A systematic review of natural language processing techniques for early detection of cognitive impairment. Mayo Clin Proc Digit Health. (2025) 3:100205. doi: 10.1016/j.mcpdig.2025.100205,
47.
ZozukNCMunkovaDKelebercovaLMunkM. Relationship between language features extracted through NLP and clinically diagnosed Alzheimer's disease and mild cognitive impairment in Slovak. Alzheimers Dement Diagn Assess Dis Monit. (2025) 17:e70122. doi: 10.1002/dad2.70122,
48.
ChouC-JChangC-TChangY-NLeeC-YChuangY-FChiuY-Let al. Screening for early Alzheimer’s disease: enhancing diagnosis with linguistic features and biomarkers. Front Aging Neurosci. (2024) 16:1451326. doi: 10.3389/fnagi.2024.1451326,
49.
HartsuikerRJBarkhuysenPN. Language production and working memory: the case of subject–verb agreement. Lang Cogn Process. (2006) 21:181–204. doi: 10.1080/01690960400002117
50.
BarbeauEJDidicMJoubertSGuedjEKoricLFelicianOet al. Extent and neural basis of semantic memory impairment in mild cognitive impairment. J Alzheimer's Dis. (2012) 30:645–57. doi: 10.3233/JAD-2011-110989
51.
HsuC-WBinderJRSebastianRThompson-SchillSLDesaiRH. Revisiting human language and speech production networks: a meta-analytic connectivity modeling study. NeuroImage. (2025) 306:121008. doi: 10.1016/j.neuroimage.2025.121008
52.
CuetosFArango-LasprillaJCUribeCValenciaCLoperaF. Linguistic changes in verbal expression: a preclinical marker of Alzheimer's disease. J Int Neuropsychol Soc. (2007) 13:433–9. doi: 10.1017/s1355617707070609
53.
PakhomovSVSChaconDWicklundMGundelJ. Computerized assessment of syntactic complexity in Alzheimer's disease: a case-study of Iris Murdoch's writing. Behav Res Methods. (2011) 43:136–44. doi: 10.1017/S1355617707070609
54.
BerishaVWangSLaCrossALissJ. Tracking discourse complexity preceding Alzheimer's disease diagnosis: a case study of president Ronald Reagan. J Alzheimer's Dis. (2015) 45:1107–13. doi: 10.3233/JAD-142763
55.
AhmedSHaighA-MFde JagerCAGarrardP. Connected speech as a marker of disease progression in autopsy-proven Alzheimer's disease. Brain. (2013) 136:3727–37. doi: 10.1093/brain/awt269,
56.
SnowdonDAGreinerLHKemperSJNanayakkaraNMortimerJA. Linguistic ability in early life and cognitive function and Alzheimer's disease in late life: findings from the Nun study. JAMA. (1996) 275:528–32. doi: 10.1001/jama.275.7.528
57.
ClarkeKMGleasonCEHohmanTJWhithamBMDowlingNM. The Nun study: key insights from 30 years of research on aging and dementia. Alzheimers Dement. (2025). doi: 10.1002/alz.14626
58.
BerishaVLissJM. Responsible development of clinical speech AI: bridging the gap between clinical research and technology. NPJ Digit Med. (2024) 7:208. doi: 10.1038/s41746-024-01199-1,
59.
VoletiRLissJMBerishaV. A review of automated speech and language features for assessment of cognitive impairment. Front Aging Neurosci. (2020) 12:588871. doi: 10.1109/JSTSP.2019.2952087
60.
de la Fuente GarciaSRitchieCWLuzS. Artificial intelligence, speech, and language processing approaches to monitoring Alzheimer's disease: a systematic review. J Alzheimer's Dis. (2020) 78:1547–74. doi: 10.3233/JAD-200888,
61.
DingKChettyMNoori HoshyarABhattacharyaT. Speech-based detection of Alzheimer's disease: a survey of methods, challenges, and clinical potential. Artif Intell Rev. (2024). 57. doi: 10.1007/s10462-024-10961-6
62.
LuzS.HaiderF.de la FuenteS.FrommD.MacWhinneyB. (2020). Alzheimer’s dementia recognition through spontaneous speech: the ADReSS challenge. In Proceedings of the Twenty-First Annual Conference of the International Speech Communication Association (Interspeech 2020) (pp. 2172–2176). International Speech Communication Association. Available online at: https://doi.org/10.21437/Interspeech.2020-2573. (Accessed March 24, 2026).
63.
LuzS.HaiderF.de la FuenteS.FrommD.MacWhinneyB. (2021). Detecting cognitive decline using speech only: the ADReSSo challenge. In Proceedings of the Twenty-Second Annual Conference of the International Speech Communication Association (Interspeech 2021) (pp. 3780–3784). International Speech Communication Association. Available online at: https://doi.org/10.21437/Interspeech.2021-1220. (Accessed March 24, 2026).
64.
BeckerJTBollerFLopezOLSaxtonJMcGonigleKL. The natural history of Alzheimer's disease: description of study cohort and accuracy of diagnosis. Arch Neurol. (1994) 51:585–94.
65.
SyedM. S. S.SyedZ. S.LechM.PirogovaE. (2020). Automated screening for Alzheimer’s dementia through spontaneous speech. In Proceedings of the Twenty-First Annual Conference of the International Speech Communication Association (Interspeech 2020) (pp. 2222–2226). Shanghai, China: International Speech Communication Association.
66.
LiJ.YuJ.YeZ.WongS.MakM.MakB.et al (2021). A comparative study of acoustic and linguistic features classification for Alzheimer's disease detection. In ICASSP 2021–2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 6423–6427). Toronto, Ontario, Canada: IEEE.
67.
Martínez-NicolásILlorenteTEMartínez-SánchezFMeilánJJG. Ten years of research on automatic voice and speech analysis of people with Alzheimer's disease and mild cognitive impairment: a systematic review article. Front Psychol. (2021) 12:620251. doi: 10.3389/fpsyg.2021.620251,
68.
XiuNVaxelaireBLiLLingZXuXHuangLet al. A study on voice measures in patients with Alzheimer’s disease. J Voice. (2025) 39:286.e13–24. doi: 10.1016/j.jvoice.2022.08.010
69.
DevlinJ.ChangM. W.LeeK.ToutanovaK. (2019). Bert: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: human Language Technologies, volume 1. Minneapolis, Minnesota. (pp. 4171–4186).
70.
Meta AI (2024) The LLaMA 3 herd of Models (Technical Report) Meta Platforms, Inc. Available online at: https://ai.meta.com/research/publications/the-llama-3-herd-of-models/ (Accessed March 24, 2026).
71.
BaevskiAZhouYMohamedAAuliM. wav2vec 2.0: a framework for self-supervised learning of speech representations. Adv Neural Inf Process Syst. (2020) 33:12449–60.
72.
QiaoY.YinX.WiechmannD.KerzE. (2021) Alzheimer’s disease detection from spontaneous speech through combining linguistic complexity and (dis)fluency features with pretrained language models. In Proceedings of the Twenty-Second Annual Conference of the International Speech Communication Association (Interspeech 2021) (pp. 3805–3809). International Speech Communication Association. Available online at: https://doi.org/10.21437/Interspeech.2021-1415 (Accessed March 24, 2026).
73.
ZhuY.ObyatA.LiangX.BatsisJ. A.RothR. M. (2021). Wavbert: exploiting semantic and non-semantic speech using wav2vec and bert for dementia detection. In Proceedings of Interspeech 2021 (Vol. 2021, p. 3790). Czech Republic: Brno.
74.
ZhuY.LinN.BalivadaK. S.HaehnD.LiangX. (2024). Adversarial text generation using large language models for dementia detection. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Miami, Florida, USA. (pp. 21918–21933).
75.
ShaoHPanYWangYZhangY. Modality fusion using auxiliary tasks for dementia detection. Comput Speech Lang. (2025) 95:101814. doi: 10.1016/j.csl.2025.101814
76.
GarcíaAMde LeonJTeeBLBlasiDEGorno-TempiniML. Speech and language markers of neurodegeneration: a call for global equity. Brain. (2023) 146:4870–9. doi: 10.1093/brain/awad253,
77.
HarperLFumagalliGGBarkhofFScheltensPO'BrienJTBouwmanFet al. MRI visual rating scales in the diagnosis of dementia: evaluation in 184 post-mortem confirmed cases. Brain. (2015) 138:2020–33. doi: 10.1093/brain/aww005,
78.
MorrisJCMohsRCRogersHFillenbaumGHeymanA. Consortium to establish a registry for Alzheimer's disease (CERAD) clinical and neuropsychological assessment of Alzheimer's disease. Neurology. (1988) 38:1445–52. doi: 10.1212/WNL.38.9.1445
79.
WechslerD. Wechsler Memory Scale—Fourth Edition (WMS-IV) Technical and Interpretive Manual. San Antonio, TX: Pearson (2009).
80.
WechslerD. Wechsler Adult Intelligence Scale—Fourth Edition (WAIS-IV) Administration and Scoring Manual. San Antonio, TX: Pearson (2008).
81.
ZimmermannPFimmB. Test of Attentional Performance (TAP), Version 2.1. Herzogenrath: Psytest (2002).
82.
HindmarchILehfeldHde JonghPErzigkeitH. The Bayer activities of daily living scale (B-ADL). Int J Geriatr Psychiatry. (1998) 13:359–66. doi: 10.1159/000051195
83.
CummingsJLMegaMGrayKRosenbergTCarusiDGornbeinJ. The neuropsychiatric inventory: comprehensive assessment of psychopathology in dementia. Neurology. (1994) 44:2308–14. doi: 10.1212/WNL.44.12.2308
84.
BeckATSteerRABrownGK. Manual for the Beck Depression Inventory–II. San Antonio, TX: Psychological Corporation (1996).
85.
YesavageJABrinkTLRoseTLLumOHuangVAdeyMet al. Development and validation of a geriatric depression screening scale: a preliminary report. J Psychiatr Res. (1983) 17:37–49. doi: 10.1016/0022-3956(82)90033-4
86.
ZigmondASSnaithRP. The hospital anxiety and depression scale. Acta Psychiatr Scand. (1983) 67:361–70. doi: 10.1111/j.1600-0447.1983.tb09716.x,
87.
GoodglassHKaplanE. The Boston Diagnostic Aphasia Examination. Philadelphia: Lea & Febinger (1983).
88.
RadfordA.KimJ. W.XuT.BrockmanG.McLeaveyC.SutskeverI. (2023). Robust speech recognition via large-scale weak supervision. In Proceedings of the Fortieth International Conference on Machine Learning (pp. 28492–28518). Honolulu, Hawaii, USA: Proceedings of Machine Learning Research.
89.
ChristiansenMHChaterN. The now-or-never bottleneck: a fundamental constraint on language. Behav Brain Sci. (2016) 39:e62. doi: 10.1017/S0140525X1500031X,
90.
FrostRBogaertsLSamuelAGMagnusonJSHoltLLChristiansenMH. Statistical learning subserves a higher purpose: novelty detection in an information foraging system. Psychol Rev. (2025). doi: 10.1037/rev0000547
91.
DeldarZGevers-MontoroCKhatibiAGhazi-SaidiL. The interaction between language and working memory: a systematic review of fMRI studies in the past two decades. AIMS Neurosci. (2020) 8:1–32. doi: 10.3934/Neuroscience.2021001,
92.
NippoldMA. Later Language Development: School-age children, Adolescents, and Young Adults. Austin, TX: PRO-ED, Inc. (2016).
93.
HuttenlocherJVasilyevaMCymermanELevineS. Language input and child syntax. Cogn Psychol. (2002) 45:337–74. doi: 10.1016/S0010-0285(02)00500-5,
94.
ThorntonRLightLL. "Language comprehension and production in normal aging". In: Handbook of the Psychology of Aging. Burlington, MA 01803, USA: Academic Press (2006). p. 261–87.
95.
YeungAIaboniARochonELavoieMSantiagoCYanchevaMet al. Correlating natural language processing and automated speech analysis with clinician assessment to quantify speech-language changes in mild cognitive impairment and Alzheimer's dementia. Alzheimer's Res Ther. (2021) 13:109. doi: 10.1186/s13195-021-00848-x,
96.
PettiUBakerSKorhonenA. A systematic literature review of automatic Alzheimer's disease detection from speech and language. J Am Med Inform Assoc. (2020) 27:1784–97. doi: 10.1093/jamia/ocaa174,
97.
UreJ. "Lexical density: a computational technique and some findings". In: CoulthardM, editor. Advances in Spoken Discourse Analysis. Leiden, The Netherlands: Routledge (1971). p. 121–35.
98.
MalvernDRichardsBChipereNDuraP. Lexical Diversity and Language Development: Quantification and Assessment. Houndmills, England: Palgrave MacMillan (2004).
99.
GuiraudP. Les caractères statistiques du vocabulaire: Essai de méthodologie. Paris, France: Presses Universitaires de France (1954).
100.
HerdanG. Type-token Mathematics: A Textbook of Mathematical Linguistics. The Hague, Netherlands: Mouton (1960).
101.
ReadJ. Assessing Vocabulary. Oxford: Oxford University Press (2000).
102.
RaynerK. Eye movements and attention in reading, scene perception, and visual search. Q J Exp Psychol. (2009) 62:1457–506. doi: 10.1080/17470210902816461,
103.
PatersonKBMcGowanVAWarringtonKLLiLLiSXieFet al. Effects of normative aging on eye movements during reading. Vision. (2020) 4:7. doi: 10.3390/vision4010007,
104.
SchroederSHyönäJLiversedgeSP. Developmental eye-tracking research in reading: introduction to the special issue. J Cogn Psychol. (2015) 27:500–10. doi: 10.1080/20445911.2015.1046877
105.
AsgariMKayeJDodgeH. Predicting mild cognitive impairment from spontaneous spoken utterances. Alzheimers Dement. (2017) 3:219–28. doi: 10.1016/j.trci.2017.01.006,
106.
FraserKCRudziczFRochonE. Using text and speech features to diagnose Alzheimer's disease. J Alzheimer's Dis. (2019) 71:1065–82. doi: 10.3233/JAD-190452
107.
DeutschP. (1996). DEFLATE Compressed Data Format Specification Version 1.3. IETF RFC 1951.
108.
ZivJLempelA. A universal algorithm for sequential data compression. IEEE Trans Inf Theory. (1977) 23:337–43. doi: 10.1109/tit.1977.1055714
109.
HuffmanDA. A method for the construction of minimum-redundancy codes. Proc IRE. (1952) 40:1098–101. doi: 10.1109/jrproc.1952.273898
110.
BoschiVCatricalàEConsonniMChesiCMoroACappaSF. Connected speech in neurodegenerative language disorders: a review. Front Psychol. (2017) 8:269. doi: 10.3389/fpsyg.2017.00269,
111.
SzatloczkiGHoffmannIVinczeVKalmanJPakaskiM. Speaking in Alzheimer's disease, is that an early sign? Importance of changes in language abilities in Alzheimer's disease. Front Aging Neurosci. (2015) 7:195. doi: 10.3389/fnagi.2015.00195,
112.
PickeringMJGarrodS. An integrated theory of language production and comprehension. Behav Brain Sci. (2013) 36:329–47. doi: 10.1017/S0140525X12001495,
113.
WillemsRMFrankSLNijhofADHagoortPvan den BoschA. Prediction during natural language comprehension. Cereb Cortex. (2016) 26:2506–16. doi: 10.1093/cercor/bhv075
114.
MartinCDBranziFMBarM. Prediction is production. Sci Rep. (2018) 8:1079. doi: 10.1038/s41598-018-19499-4
115.
RyskinRANieuwlandMS. Prediction during language comprehension: what is next?Trends Cogn Sci. (2023) 27:1032–52. doi: 10.1016/j.tics.2023.08.003,
116.
FedermeierKD. Thinking ahead: the role and roots of prediction in language comprehension. Psychophysiology. (2007) 44:491–505. doi: 10.1111/j.1469-8986.2007.00531.x,
117.
ArnonISniderN. More than words: frequency effects for multi-word phrases. J Mem Lang. (2010) 62:67–82. doi: 10.1016/j.jml.2009.09.005
118.
AmbridgeBKiddERowlandCFTheakstonAL. The ubiquity of frequency effects in first language acquisition. J Child Lang. (2015) 42:239–73. doi: 10.1017/S030500091400049X,
119.
WiechmannD.QiaoY.KerzE.MatternJ. (2022). Measuring the impact of (psycho-)linguistic and readability features on eye movement patterns. In Proceedings of ACL 2022 (pp. 5276–5290). Dublin, Ireland: ACL.
120.
KerzE.HeilmannA.NeumannS. (2019). L2 processing advantages of multiword sequences: evidence from eye-tracking. In Proceedings of the Joint Workshop on Multiword Expressions and WordNet (pp. 60–69). Florence, Italy: ACL.
121.
HernándezMCostaAArnonI. More than words: multiword frequency effects in non-native speakers. Lang Cogn Neurosci. (2016) 31:785–800. doi: 10.1080/23273798.2016.1152389
122.
KerzEWiechmannD. "Individual differences in L2 processing of multi-word phrases: effects of working memory and personality". In: MitkovR, editor. Computational and Corpus-Based Phraseology. London, UK: Springer (2017). p. 325–40.
123.
DingJDuJWangHXiaoS. A novel two-stage feature selection method based on random forest and improved genetic algorithm for enhancing classification in machine learning. Sci Rep. (2025) 15:16828. doi: 10.1038/s41598-025-01761-1,
124.
SuiQLiGPengYZhangJZhangYZhaoR. Scalable and robust machine learning framework for HIV classification using clinical and laboratory data. Sci Rep. (2025) 15:18727. doi: 10.1038/s41598-025-00085-4,
125.
PedregosaFVaroquauxGGramfortAMichelVThirionBGriselOet al. Scikit-learn: machine learning in Python. J Mach Learn Res. (2011) 12:2825–30.
126.
ChenT.GuestrinC. (2016). XGBoost: a scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). ACM.
127.
LundbergSMLeeS-I. "A unified approach to interpreting model predictions". In: GuyonIet al, editors. Advances in Neural Information Processing Systems 30. Red Hook, New York, USA: Curran Associates, Inc. (2017). p. 4765–74.
128.
LundbergSMErionGChenHDeGraveAPrutkinJMNairBet al. From local explanations to global understanding with explainable AI for trees. Nat Mach Intell. (2020) 2:252–9. doi: 10.1038/s42256-019-0138-9
129.
LundbergS.SHAP (SHapley Additive exPlanations). Python Package Version v0.50.0. (2025). Available online at: https://github.com/shap/shap. (Accessed March 24, 2026).
130.
GarrigaRMasJAbrahaSNolanJHarrisonOTadrosGet al. Machine learning model to predict mental health crises from electronic health records. Nat Med. (2022) 28:1240–8.
131.
MahajanPBathsV. Acoustic and language based deep learning approaches for Alzheimer's dementia detection from spontaneous speech. Front Aging Neurosci. (2021) 13:623607. doi: 10.3389/fnagi.2021.623607
132.
YangQLiXDingXXuFLingZ. Deep learning–based speech analysis for Alzheimer’s disease detection: a literature review. Alzheimer's Res Ther. (2022) 14:186. doi: 10.1186/s13195-022-01131-3,
133.
ShiMCheungGShahamiriSR. Speech and language processing with deep learning for dementia diagnosis: a systematic review. Psychiatry Res. (2023) 329:115538. doi: 10.1016/j.psychres.2023.115538,
134.
MobtahejPGouronSTJavaheriRLeDNEl-KhouryANaingTHet al. Transformer-based deep learning approaches for speech-based dementia detection: a systematic review. IEEE J Biomed Health Inform. (2025) 30:2034–48. doi: 10.1109/JBHI.2025.3595781,
135.
SuHLiXBornSHoneyCJChenJLeeHet al. Neural dynamics of spontaneous memory recall and future thinking in the continuous flow of thoughts. Nat Commun. (2025) 16:6433. doi: 10.1038/s41467-025-61807-w
136.
HsuC-Wet al. Revisiting human language and speech production network: a meta-analytic connectivity modeling study. NeuroImage. (2025) 306:121008. doi: 10.1016/j.neuroimage.2025.121008,
137.
SperlingRADickersonBCPihlajamakiMVanniniPLaViolettePSVitoloOVet al. Functional alterations in memory networks in early Alzheimer's disease. Ann Neurol. (2009) 66:389–98. doi: 10.1007/s12017-009-8109-7
138.
BucknerRLSepulcreJTalukdarTKrienenFMLiuHHeddenTet al. Cortical hubs revealed by intrinsic functional connectivity: mapping, assessment of stability, and relation to Alzheimer's disease. J Neurosci. (2009) 29:1860–73. doi: 10.1523/JNEUROSCI.5062-08.2009,
139.
AgostaFGalantucciSFilippiM. Advanced magnetic resonance imaging of neurodegenerative diseases. Brain Connect. (2015) 5:307–33. doi: 10.1007/s10072-016-2764-x
140.
PereiraJBMijalkovMKakaeiEMecocciPVellasBTsolakiMet al. Disrupted network topology in patients with stable and progressive mild cognitive impairment and Alzheimer's disease. Cereb Cortex. (2018) 28:1993–2006. doi: 10.1093/cercor/bhw128
141.
ChoSDoughertyCCPetrellaJRSheldonFCDoraiswamyPM. Lexical and acoustic speech features relating to Alzheimer disease pathology. Neurology. (2022) 99:e2417–27. doi: 10.1212/WNL.0000000000200581,
142.
HajjarIBrownDLFairchildJKMengYParraCSanchezDet al. Development of digital voice biomarkers and associations with cognition, cerebrospinal biomarkers, and neural representation in early Alzheimer's disease. Alzheimers Dement. (2023) 15:e12382. doi: 10.1002/dad2.12393
143.
García-GutiérrezJBenítez-AndoneguiAZangYGonzález-OrtegaGElsheikhSEcay-TorresMet al. Harnessing acoustic speech parameters to decipher amyloid status in mild cognitive impairment. Front Neurosci. (2023) 17:1221401. doi: 10.3389/fnins.2023.1221401
144.
GabirondoGMuñoz-RuizMEcay-TorresMGonzález-OrtegaGSebastiánMVToledoMet al. Speech biomarkers predict amyloid status in cognitively unimpaired older adults. Intell Based Med. (2025) 9:100110. doi: 10.1016/j.ibmed.2025.100306
145.
FarzanaF.DebnathB.HsuC.-W.RobinJ.VogelA. P. (2024). A spoken language corpus for early Alzheimer's disease detection with cerebrospinal fluid and neuroimaging biomarkers (SLaCAD). In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 14562–14575). Torino, Italy: European Language Resources Association.
146.
WangYLiXZhangHZhouJLiuYLiJet al. Speech digital biomarker combined with fluid biomarkers improves prediction of cognitive impairment in memory clinic patients. Alzheimer's Res Ther. (2025) 17:25. doi: 10.1186/s13195-025-01877-6
147.
YingYYangTZhouH. Multimodal fusion for Alzheimer's disease recognition. Appl Intell. (2023) 53:16029–40. doi: 10.1007/s10489-022-04255-z
148.
FuYXuLZhangYZhangLZhangPCaoLet al. Classification and diagnosis model for Alzheimer's disease based on multimodal data fusion. Medicine. (2024) 103:e41016. doi: 10.1097/MD.0000000000041016,
149.
LinKWashingtonPY. Multimodal deep learning for dementia classification using text and audio. Sci Rep. (2024) 14:13887. doi: 10.1038/s41598-024-64438-1,
150.
AgmonGPradhanSAshSNevlerNLibermanMGrossmanMet al. Automated measures of syntactic complexity in natural speech production: older and younger adults as a case study. J Speech Lang Hear Res. (2024) 67:545–61. doi: 10.1044/2023_JSLHR-23-00009,
151.
CalzàLGagliardiGRossini FavrettiRTamburiniF. Linguistic features and automatic classifiers for identifying mild cognitive impairment and dementia. Comput Speech Lang. (2021) 65:101113. doi: 10.1016/j.csl.2020.101113
152.
EyigozEMathurSSantamariaMCecchiGNaylorM. Linguistic markers predict onset of Alzheimer's disease. EClinicalMedicine. (2020) 28:100583. doi: 10.1016/j.eclinm.2020.100583,
153.
AgmonGChoSAshSCousinsKAQBlennowKZetterbergHet al. Automatic quantification of syntactic complexity in natural spontaneous speech of people with primary progressive aphasia. Aphasiology. (2026) 40:561–82. doi: 10.1080/02687038.2025.2462282
154.
KemperSThompsonMMarquisJ. Longitudinal change in language production: effects of aging and dementia on grammatical complexity and propositional content. Psychol Aging. (2001) 16:600–14. doi: 10.1037/0882-7974.16.4.600
155.
Bartrés-FazDDemnitz-KingHCabello-ToscanoMVaqué-AlcázarLSaundersRTouronEet al. Psychological profiles associated with mental, cognitive and brain health in middle-aged and older adults. Nat Ment Health. (2025) 3:92–103. doi: 10.1038/s44220-024-00361-8
156.
AschwandenDGerstorfDTerraccianoASpiroAWagnerGGHoppmannCA. Is personality associated with dementia risk? A meta-analytic investigation. Ageing Res Rev. (2021) 67:101269. doi: 10.1016/j.arr.2021.101269,
157.
TerraccianoAStephanYLuchettiMSutinARWilsonRS. Is neuroticism differentially associated with risk of Alzheimer’s disease, vascular dementia, and frontotemporal dementia?J Psychiatr Res. (2021) 138:34–40. doi: 10.1016/j.jpsychires.2021.03.039,
158.
KerzEZanwarSQiaoYWiechmannD. Toward explainable artificial intelligence for mental health detection based on language behavior. Front Psych. (2023) 14:1219479. doi: 10.3389/fpsyt.2023.1219479,
159.
UNESCO. (2012). International Standard Classification of Education: ISCED 2011. UNESCO Institute for Statistics. Available online at: https://www.uis.unesco.org/en/methods-and-tools/isced (Accessed March 24, 2026).
Summary
Keywords
Alzheimer’s disease, CSF biomarkers, digital biomarkers, explainable AI, machine learning, natural language processing
Citation
Wiechmann D, Kerz E, Albrecht M, Qiao Y, Pinho J, Reetz K and Costa AS (2026) Detecting CSF-validated Alzheimer’s disease from spontaneous speech in German: an interpretable end-to-end machine-learning framework. Front. Neurol. 17:1780783. doi: 10.3389/fneur.2026.1780783
Received
04 January 2026
Revised
13 March 2026
Accepted
16 March 2026
Published
09 April 2026
Volume
17 - 2026
Edited by
Pedro Gomez-Vilda, Neuromorphic Speech Processing Laboratory, Spain
Reviewed by
Jesús B. Alonso-Hernández, University of Las Palmas de Gran Canaria, Spain
Galit Agmon, Bar-Ilan University, Israel
Updates
Copyright
© 2026 Wiechmann, Kerz, Albrecht, Qiao, Pinho, Reetz and Costa.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Daniel Wiechmann, d.wiechmann@uva.nl; Milena Albrecht, malbrecht@ukaachen.de; Ana Sofia Costa, acosta@ukaachen.de
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.