ORIGINAL RESEARCH article

Front. Lang. Sci., 07 September 2026

Sec. Bilingualism

Volume 5 - 2026 | https://doi.org/10.3389/flang.2026.1929126

Convergence between AI chatbot-calculated microstructure measures and SALT analysis of child language samples

  • Department of Speech, Language, and Hearing Sciences, University of Colorado, Boulder, CO, United States

Abstract

This study examined the association and agreement between microstructure calculations from four large language model (LLM) chatbots and those generated by Systematic Analysis of Language Transcripts (SALT) software. Eighty-four typically developing children (35 girls; ages 2–9) participated, with 30 Japanese-English bilinguals and 54 English monolinguals. Narrative language samples were collected and coded with SALT conventions. Transcripts were randomly ordered and analyzed using two prompts. Using a general prompt, transcripts were entered into each LLM chatbot (ChatGPT, Copilot, MagicSchool, and Gemini), microstructure outputs were recorded, then the process was repeated with the same order using the more specific prompt. Results showed that across LLM chatbots, measures requiring relatively simple word or morpheme counts (MLUw, MLUm, NTW, NDW) generally exhibited stronger associations with SALT than the subordination index (SI), which showed greater variability. While intraclass correlation coefficients within measures were more variable, association and agreement coefficients were generally consistent with each other. Children's age and language backgrounds were associated with differences in the magnitude of some LLM–SALT coefficients. Finally, prompt specificity did not strengthen microstructure calculations for any of the LLM chatbots. Findings provide preliminary methodological evidence that some SALT-calculated microstructure measures may be more amenable to LLM chatbot-assisted calculation than others, with substantial associations and generally strong agreement observed for several measures; however, due to the exploratory nature of the current study, their clinical utility, decision-level accuracy, and workflow implications require further investigation.

1 Introduction

Speech-language pathologists (SLPs) use language sample analysis (LSA) to assess clients' natural language by analyzing features such as morphosyntax, vocabulary, semantics, discourse, and narrative skills (; Leadholm and Miller, 1994). Unlike standardized assessments, LSA captures how individuals integrate language skills in functional contexts and can identify communication difficulties that standardized tests may miss (Leadholm and Miller, 1994). LSA is recognized as a key component of comprehensive language assessment, demonstrating acceptable diagnostic accuracy (Ramos et al., 2022), associations with language and literacy outcomes (), and endorsement in the American Speech-Language-Hearing Association's Preferred Practice Patterns (), particularly for school-age children and those from linguistically diverse backgrounds. Language samples provide an ecologically valid index of children's functional language and are especially informative when used alongside standardized assessments (; ; Justice et al., 2006).

Despite these benefits, LSA is used less frequently than recommended because of time demands and large caseloads (Kemp and Klee, 1997; Pavelko et al., 2016; Wilder and Redmond, 2024). This gap contributes to greater reliance on standardized assessments and increased risk of bias when assessing multilingual children (Justice et al., 2006; Kapantzoglou et al., 2017). In response, school-based SLPs have begun exploring generative AI tools, including large language model (LLM) chatbots—AI systems trained on internet-scale data to interact conversationally () – despite limited evidence supporting their validity for assessment purposes (; Lee and Yoon, 2025). The present study examines the association and agreement between LLM chatbot-calculated and microstructure measures calculated by Systematic Analysis of Language Transcripts (SALT; Miller et al., 2019) as an initial step toward determining whether LLM chatbots can accurately calculate established language sample microstructure measures under controlled conditions. Association reflects the strength of the relationship between LLM- and SALT-calculated measures, whereas agreement reflects the extent to which the two methods yield similar values.

With the advent of LLM chatbots, such as ChatGPT, there has been a relatively recent surge in the use of artificial intelligence (AI) in education (Kohnke et al., 2023; Pokrivcakova, 2019; Rahman et al., 2024; Zhang et al., 2024b) and communication sciences and disorders (CSD; ; ; ; ; ; ; Lee and Yoon, 2025; Villamil et al., 2020; Zhang et al., 2024a). However, LLM chatbots have primarily been used to support ancillary clinical tasks, such as creating educational materials () or clinical note writing (). Whether these tools can be relied upon to analyze and compute core microstructure measures from child language samples with sufficient accuracy, however, remains an open empirical question. The current study examines this question using four LLM chatbots: ChatGPT, Copilot, MagicSchool, and Gemini. To our knowledge, no prior published studies have examined the association or agreement between LLM chatbot-calculated and established microstructure measures in LSA. By empirically assessing whether LLM chatbots can calculate accurate microstructure measures—such as mean length of utterance in words or morphemes (MLUw, MLUm), number of different words (NDW), number of total words (NTW), subordination index (SI), and mazes—this study provides evidence that LLM chatbot-calculated microstructure measures can show substantial correspondence with SALT-calculated measures under controlled conditions, providing a foundation for future research on AI-assisted LSA.

1.1 The clinical utility of language sample analysis

Language sample analysis is a core component of language assessment. Unlike norm-referenced standardized assessments, LSA is criterion-referenced, with the goal of describing a child's performance within a target domain (). In LSA, performance is typically evaluated through microstructures—the internal linguistic elements of a narrative, such as conjunctions, noun phrases, and clauses (Justice et al., 2006). Microstructures are commonly grouped into productivity (amount of language produced), lexical diversity (number of unique words), and syntactic complexity, measured through indices such as MLUw or more advanced clausal measures like the SI (Košutar et al., 2022). Microstructures can be calculated by an annotation system, two common ones include the Codes for the Human Analysis of Transcripts (MacWhinney, 2000; MacWhinney and Fromm, 2022) and SALT (Miller et al., 2019). However, Pavelko et al. (2016) surveyed SLPs and found that many used a self-designed LSA protocol (46%), followed by SALT (24%) being the second most used.

Prior studies have examined microstructure measures for indexing linguistic development (Miller and Chapman, 1981) and differentiating typically developing (TD) from language-impaired children (Rice et al., 2010). Microstructures are also valuable for assessing linguistically diverse children. Parker and Brorson (2005) proposed MLUw as an alternative to MLUm due to its near-perfect correlation with MLUm (r = 0.998), reduced calculation time, and greater stability across languages and dialects. Supporting its validity, Klee et al. (2004) found that Cantonese children's MLUw and lexical diversity correlated with age (27–68 months) and differentiated TD from a specific language impairment (SLI) with 98% accuracy. Lexical diversity is likewise correlated with age and serves as a reliable developmental index (), with subtle differences observed between TD and SLI children (Owen and Leonard, 2002). For syntactic complexity, significant group differences in SI scores have been reported in Spanish–English bilingual children with and without a primary language impairment (Kapantzoglou et al., 2017). Additionally, microstructure measures can uncover differences between bilingual and monolingual children within the same language. For example, Otwinowska et al. (2020) found differences in Polish-English bilinguals' MLU compared to Polish monolinguals' due to unnecessary pronoun use. found lower NDW in Japanese-English bilinguals' English narratives compared to English monolinguals. Finally, in a meta-analysis, Ramos et al. (2022) found consistent acceptable sensitivity and specificity for measures of grammaticality, or composite measures, which included either MLU or a grammaticality measure. Overall, these studies point to LSA being clinically valuable across populations, particularly given the bias of standardized assessments (Justice et al., 2006).

1.2 Artificial intelligence in education, clinical practice, and research

AI has been discussed in the field of CSD since 1985, with a rapid increase in related publications beginning around 2012 (Zhang et al., 2024a). This growth is partly attributed to expanded language sample databases, such as the Child Language Data Exchange System (CHILDES; MacWhinney, 2000), which support machine learning approaches well (). The resulting surge in AI research has supported applications that enhance clinical decision making (; ; ; Ramesh and Assaf, 2021), automate administrative tasks and streamline workflows (; ; Lee and Yoon, 2025; Liu et al., 2023), and improve speech therapy outcomes related to speech or language (; Bhardwaj et al., 2025; ).

In their bibliometric analysis of AI in CSD, Zhang et al. (2024a) found that the main areas of research implementing AI into CSD included autism, dysarthria, dementia, Parkinson's disease, and aphasia. Clinicians, however, are using LLM chatbots like ChatGPT. LLM chatbots, like ChatGPT, are language models that use machine learning techniques to produce human-like text. They are trained on a large corpus of text, with 300 billion words, as input to a neural network, which is a mathematical model that resembles how the human brain works, composed of 175 billion parameters (). LLM chatbots are user-friendly and can be used for patient engagement, session planning, medical education, and clinical decision support. For example, used ChatGPT to assist in therapy to help a patient with aphasia achieve their writing goals. Kohnke et al. (2023) implemented ChatGPT into language education as a stress-free conversation partner that can provide feedback and adjust the complexity of its responses. introduced an LLM chatbot-powered ambient listening device to improve medical documentation and maximize the time spent with the patient; they found that time was saved, the clinician was able to focus more on the client rather than taking notes, and the quality of the writing improved. combined AI with clinician-led therapy for /ɹ/ productions. The five participants in their study showed improvement in /ɹ/ articulation in response to an AI-assisted treatment protocol. Finally, evaluated the reliability and accuracy of ChatGPT for calculating various microstructure measures from children's narrative samples. They found that GPT-4o performed most accurately with simplest tasks such as counting words, but accuracy dropped with more complex tasks such as calculating MLU in words. In general, LLM chatbots have been explored as tools to support in clinical practice. However, research on AI in CSD shows budding potential for its use in helping SLPs with a variety of tasks from identification and prognostics to treatment and documentation.

Beyond CSD, researchers have applied LLMs to qualitative text analysis (Jenner et al., 2025; Jones et al., 2019; Tai et al., 2024), text annotation (), and educational assessment scoring (Organisciak and Acar, 2026). Jones et al. (2019) found that (Bidirectional Encoder Representations from Transformers) BERT predicted children's narrative macrostructure scores with accuracy comparable to human raters. Jenner et al. (2025) reported that both Claude 3 Opus and GPT-o1 substantially reduced analysis time, with Claude producing results that more closely matched human analyses. Similarly, Tai et al. (2024) found that ChatGPT 3.5 could support deductive qualitative analysis when provided with a codebook and effective prompts, though some errors remained. Organisciak and Acar (2026) demonstrated that prompting techniques such as self-confidence estimation, weighted scoring, and ensemble models improved LLM scoring of the Alternate Uses Task. However, they noted that LLMs are better suited for qualitative than quantitative scoring, arguing that by asking them to perform quantitative scoring, “we are, in essence, asking them to solve equations with a typewriter rather than a calculator” (Organisciak and Acar, 2026). Despite these advances, LLMs still have important limitations.

A common problem with LLM chatbots is “hallucinations,” which are inaccuracies and made-up information produced by LLM chatbots, due to the output relying on statistical patterns and probabilities, rather than verifying the truthfulness of the output (; Hicks et al., 2024). Hicks et al. (2024) argue that, because LLM chatbots cannot be concerned with truth, and because they are specifically designed to produce text that looks like the truth, their output could be deemed, in their words, “bullshit.” That LLM chatbots can generate lies and do so confidently is a major concern. Kohnke et al. (2023) pointed out that ChatGPT sounds definitive, uses no hedging, and thus can convince the user that it is correct and accurate, even when it is completely wrong. This is an issue for individuals who cannot effectively evaluate the truth of these LLM chatbots' claims, such as children, second language learners, individuals with cognitive disabilities, or even individuals with minimal knowledge about AI. Furthermore, LLM chatbots are specifically engineered to produce output that makes them seem like they are human. It is important to note that ChatGPT and similar LLM chatbots do not “understand” language as humans do. cautions that assuming AI understanding is risky, as these systems only simulate comprehension. Their opaque internal processes—the AI “black box” problem—are often poorly understood by users (). While AI holds significant potential for clinical SLP practice, LLM chatbots can also propagate misinformation and bias (). For example, Lewis et al. (2025) found that AI-generated images perpetuated cultural, linguistic, and disability biases, highlighting the risk of reinforcing racism and ableism.

In addition to ethical concerns, there is a serious concern for data privacy. Ramesh and Assaf (2021) point out that speech data and linguistic patterns, like fingerprints, are unique. AI and related tools are data-driven, and truly informed consent requires that the individual is aware of and understands what they are consenting to (). When a clinician puts the linguistic data of a child into an LLM chatbot, do they know if it is a secure place for a client's data? Even if it is deidentified, is it still safe to put into an LLM chatbot? For example, in the case of AI use in clinical practice, point out that a family or client may consent to the use of AI in clinical practice while also not consenting to the speech or language data being reused by the AI company to improve their products. ChatGPT has an option called “Improve the model for everyone,” which can be turned off. In doing this, conversations with ChatGPT will still be saved in the chat history, but they will not be used to train ChatGPT. Clinicians must be aware of this fact before implementing LLM chatbots into their practice to protect client data. For these reasons and more, data privacy pertaining to the implementation of AI into CSD is an often-cited concern (; ; ; ; ; ; Zhang et al., 2024a). argue that higher education institutions should prepare future SLPs to critically and ethically evaluate LLM chatbots, because despite the surge in research pertaining to AI, Villamil et al. (2020) found significant gaps in academic literature on AI pertaining to its use for audiologists and SLPs. This is especially true for school-based clinicians. AI in CSD is not inevitable; it is currently happening despite limited empirical evidence. Thus, continued investigation into the accuracy, agreement, and limitations of LLM chatbot-calculated language sample measures is both timely and important.

1.3 The current study

The current study examined the association and agreement between microstructure measures calculated by LLM chatbots and the same microstructure measures calculated using SALT as a criterion measure (Mislevy and Rupp, 2010). Evaluations of new measurement methods commonly begin by examining their association with an established reference, although strong associations alone do not demonstrate measurement equivalence (; Mislevy and Rupp, 2010). In CSD, similar approaches have been used to compare caregiver questionnaires with behavioral measures of vocabulary, syntax, and prelinguistic skills (; Marchman and Martínez-Sussmann, 2002; Thal et al., 1999; Wetherby et al., 2002). Accordingly, the present study examined both association, to evaluate the strength of the relationship between LLM chatbot- and SALT-calculated measures, and agreement, to evaluate the extent to which the two methods produced similar values (Shrout and Fleiss, 1979). This investigation serves as a preliminary step toward evaluating the potential of LLM-assisted language sample analysis. Currently, AI-assisted methods that automate transcribing and coding child language samples for LSA are limited (Lammert et al., 2025; Lüdtke et al., 2023). Time constraints appear to be the main barrier for SLPs conducting LSA (Pavelko et al., 2016; Wilder and Redmond, 2024), even after providing SLPs with comprehensive LSA training (Klatte et al., 2022). This study provides preliminary methodological evidence to inform future research on clinical applications of AI-assisted LSA. Specific research questions are:

  • To what extent do microstructure measures (i.e., MLUm, MLUw, SI, NTW, NDW & mazes) calculated by LLM chatbots associate and agree with those calculated by SALT?

  • What are the patterns of association and agreement across the different microstructure measures?

  • Does age group (preschool vs. school-age) or language group (bilingual vs. monolingual) affect the magnitude of the association between LLM chatbot-calculated and SALT-calculated microstructure measures?

  • Does greater prompt specificity alter the association between LLM chatbot-calculated and SALT-calculated MLUm & SI?

For RQ1 and RQ2, we hypothesize that the LLM chatbot-calculated microstructure measures will be associated with and agree with the corresponding SALT microstructure measures but based on findings regarding LLM chatbot math abilities (; ; King, 2023), and the findings of , the calculations for utterance-level measures like SI and MLUm may be less strongly associated than simpler, word-level calculations like NTW. For RQ3, we hypothesize that LLM chatbot-calculated microstructure measures may vary by age group, given systematic differences in transcript length and syntactic complexity between preschool and school-age narratives. We also test whether language group moderates these associations, recognizing that bilingual and monolingual children may differ in lexical and morphosyntactic patterns even within English-only samples (e.g., ; Otwinowska et al., 2020). For RQ4, we hypothesize that greater prompt specificity will improve the output of the LLM chatbots, leading to stronger microstructure calculation associations with SALT (; Kirilenko, 2026; Oppenlaender et al., 2025).

2 Methods

This study was approved by the University of Colorado Boulder Institutional Review Board (Protocol number: 23-0739). Informed consent was collected from all participants.

2.1 Participants

Participants were 84 typically developing children (35 girls) ages 2 to 9 (M = 75.62 months, SD = 22.51 months); 54 were English monolinguals ages 3 to 9 (M = 78.3 months, SD = 22.47), and 30 were Japanese-English bilinguals ages 2 to 8 (M = 70.8 months, SD = 22.14). The participants were drawn from a larger study on narrative language development (). Aside from the two-year-olds, all participants attend English public schools near the Denver metropolitan area and reported no speech or language concerns based on a caregiver questionnaire. Only the English language samples were used.

2.2 Language sampling procedures

To elicit language samples, MAIN () was used. MAIN was developed to assess narrative production and comprehension in children aged 2 to 10 years old. In the current study, we examined the narratives of children younger than the MAIN age range. However, other studies have utilized MAIN for participants outside of the age range (e.g., ). It can be used as a story-tell, retell, or model story. The MAIN Dog and Cat parallel stories were used as a story retell in this study. See Table 1 for microstructure information.

Table 1

SALT Microstructure Data
GroupMLUwMLUmSINTWNDWMazes
School-age (n = 46)
M7.668.251.18108.3349.895.17
SD1.291.430.1329.2311.793.34
Range5.56–11.155.88–12.20.86–1.4241–16424–740–15
Preschool (n = 38)
M5.616.151.0484.2639.474.39
SD1.561.620.1428.6111.844.77
Range2.26–8.52.57–90.67–1.3641–16420–640–19
MLUwMLUmSINTWNDWMazes
Bilingual (n = 30)
M6.296.831.0784.3338.533.27
SD2.162.280.1734.6113.92.59
Range2.26-11.152.57-120.67-1.3841-15520-710-9
Monolingual (n = 54)
M6.987.571.14104.7248.875.69
SD1.421.510.1326.7910.674.45
Range3.58–10.84.12–12.20.86–1.4257–16429–740–19

Participant demographic & SALT microstructure data.

MLUw, mean length of utterance in words; MLUm, mean length of utterance in morphemes; SI, subordination index; NTW, number of total words; NDW, number of different words.

Participants were given the option of three stories in envelopes to choose from; however, each envelope contained the same story, which was to control for the effect of shared knowledge (). After the child chose the envelope, the examiner displayed the pictures while reading the story, then prompted the child to tell the story back, “in their own words.” All participants' narratives were recorded using OM System Olympus WS-882 Digital Recorders.

2.3 Transcription

The children's narratives were transcribed and analyzed using SALT software, a program that was designed for making LSA easy and accessible for researchers and clinicians (Version 20; Miller et al., 2019). The audio recordings were transcribed and coded following SALT conventions by the PI and a trained RA. The transcripts were coded to mark individual words and bound morphemes to calculate MLUm, MLUw, NDW, and NTW. To calculate SI, the number of clauses was marked with the code “SI-[number]” to indicate the number of clauses per utterance (e.g., SI-1 for one clause, SI-2 for a main and subordinate clause, etc.). Finally, mazes—filled pauses, false starts, and repetitions—were marked in parentheses [e.g., “And (then um) then (h*) he left.”]. Comparison of 20 transcripts coded by both the RA and the PI showed a 94.9% item-by-item agreement on SALT codes.

To elicit the same microstructure measures from each AI chatbot, following transcription of the language sample, the transcript was separated by communication units (C-units). C-units are defined as an independent clause with its modifiers, which includes one main clause and all subordinate clauses attached to it (Miller et al., 2019). Individual C-units were put on their own line. Coordinating conjunctions, such as, “and,” “but,” “so,” and “then,” mark main clauses, and thus indicate a new C-unit. Subordinating conjunctions, such as, “because,” that,” when,” “who,” “after,” “before,” “so (that),” “which,” “although,” “if,” “unless,” “while,” “as,” “how,” “until,” “as__as,” “like,” “where,” “since,” “although,” “who,” “before,” and “how,” mark subordinate clauses, and thus are kept on the same line of the transcript as the main clause. In addition to C-units, mazes were put into parentheses. Aside from separating lines by C-units and marking mazes, nothing else was done to prepare the transcripts for the LLM chatbots. See Appendix 1 for an example transcript.

2.4 LLMs & AI prompts

Four LLMs were used for this study: OpenAI's ChatGPT, Microsoft's Copilot, MagicSchool AI, and Google Gemini. These four LLMs were chosen because they have a free version, making them easily accessible to SLPs, and they are included as part of common email platforms that school district use (i.e., Gmail and Outlook), they are specifically designed for educators (i.e., MagicSchool), or they are very widely used (i.e., ChatGPT; see ).

Two AI prompts were used (see Appendix 2). The first prompt was shorter and less detailed. It was designed to be as simple and clear as possible, and to account for tokenizer limits (), and the second prompt included the same verbiage as the first prompt, but added more specificity and calculation parameters for MLUm and SI. Research shows that AI chatbots struggle with calculations and can display bias in numerical responses (; ; King, 2023), but “prompt engineering,” or adding more details and context to the prompt, can improve AI chatbot output (; Oppenlaender et al., 2025).

The first prompt was broad and simply asked to provide the requested microstructure measures for the transcript provided at the end of the prompt, after the colon. The only restriction added was regarding mazes, which was “the items in the parentheses.”

The second prompt was identical to the first prompt, with extra specifications provided in parentheses about calculating MLUm and SI. This information was taken from the SALT Standard Convention manual (Miller et al., 2019).

2.5 Procedure

The transcripts were randomly ordered from one to 84 then kept consistent for both prompts. Beginning with prompt one, the first transcript was added to the prompt after a colon, then the whole body of text was inputted into all four LLM chatbots starting with ChatGPT, then Copilot, then MagicSchool, and finally Gemini. The outputted microstructure measures were noted, then this process was repeated for the second transcript and so on until all 84 transcripts had been analyzed. This process was repeated with prompt two. Data collection was done in September, October, and November of 2025 across several days for both prompts, with a data collection session typically being 10 to 15 prompts to each LLM chatbot. See Appendix 3 for an example ChatGPT response.

The free version of each LLM chatbot was used and kept on the standard settings provided when you visit their webpages. For ChatGPT, this was “ChatGPT (GPT-5),” not “ChatGPT Plus.” For Copilot, this was “Smart (GPT-5),” not “Quick response,” “Think deeper,” “Study and learn,” or “Search.” For MagicSchool, this was “Raina (Chatbot)” set on the “Fastest” setting, not “Smartest” or “Thinking.” Finally, for Gemini, this was the “Fast” setting, not “Thinking with 3 Pro.”

No additional prompts were given to the LLM chatbots, except if the microstructure measures were coming out clearly wrong due to the incorrect calculation method. For example, MagicSchool, Copilot, and Gemini each, on separate occasions, one time at the beginning of a data collection session, calculated SI based on the number of subordinate clauses divided by the number of utterances, resulting in abnormally small SI calculations. SALT calculates SI based on the total number of clauses divided by the number of utterances (Miller et al., 2019). When this was the case, the LLM chatbots were immediately given the follow-up prompt: “SI is calculated by dividing the total number of clauses by the total number of utterances,” after which, each of the LLM chatbots would correct their calculations. The calculations given based on the follow-up prompt were used in the current study to examine the LLM chatbot assisted microstructure measure calculations based on the calculation method that SALT uses.

3 Data analysis

To address RQ1 & RQ2, two-way random effects, absolute agreement ICCs were calculated to examine agreement, and linear regression analyses were conducted using each LLM chatbot-calculated microstructure measure and its corresponding SALT-calculated measure to examine the strength of the associations. Standardized β coefficients and ICCs were used to describe the patterns of association and agreement across LLM chatbots and SALT-calculated microstructure measures.

To answer RQ3, forward regression analyses were conducted, with the corresponding SALT-calculated measure as the dependent variable. An unconditional model was first fitted, followed by the addition of the dichotomous variables age group (preschool & school-age) and language group (bilingual & monolingual), and the interaction terms between the LLM chatbot-calculated microstructure measure and age group and language group. The interaction terms were the primary effects of interest because they tested whether the association between the LLM chatbot-calculated and SALT-calculated measures differed by age group or language group.

To address RQ4, separate multiple regression models were conducted for MLUm and SI for each LLM chatbot, with the corresponding SALT-calculated microstructure measure as the dependent variable. Independent variables included the LLM chatbot-calculated microstructure measure (continuous), prompt number (Prompt 1 vs. Prompt 2; dichotomous), and their interaction. The interaction term was the primary effect of interest because it tested whether prompt specificity modified the association between the LLM chatbot-calculated and SALT-calculated measures.

4 Results

4.1 Convergence between LLM chatbot-calculated and SALT microstructure measures

Descriptive statistics of SALT and LLM chatbot microstructure calculations are presented in Table 2. For RQ1 & RQ2, ICCs are included in Table 3 and standardized β coefficients in Table 4. Additionally, “Model 1” in Tables 5 through 10 shows that the LLM chatbot-calculated microstructure measures were significantly associated with the corresponding SALT microstructure measure (ps < 0.001).

Table 2

SALT or LLM ChatbotMLUwMLUmMLUm (2)SISI (2)NTWNDWMazes
SALT6.737.31.1197.4445.184.82
ChatGPT6.587.377.531.150.9595.3349.544.73
Copilot6.98.177.631.241.0198.743.515.26
MagicSchool6.527.388.141.180.9894.0645.335.02
Gemini5.896.666.371.111.0483.3639.985.64

Mean microstructure calculations for SALT and each LLM chatbot.

MLUw, mean length of utterance in words; MLUm, mean length of utterance in morphemes; SI, subordination index; NTW, number of total words; NDW, number of different words. “(2)” is the calculation mean from the second, more specific prompt. There was no second calculation for SALT.

Table 3

MeasureChatGPTCopilotMagic SchoolGemini
MLUw0.930.920.970.81
MLUm0.910.910.930.82
NTW0.940.940.970.76
NDW0.740.730.970.81
SI0.450.510.390.65
Mazes0.930.920.890.91

Intraclass correlation coefficients by measure and LLM chatbot.

MLUw, mean length of utterance in words; MLUm, mean length of utterance in morphemes; SI, subordination index; NTW, number of total words; NDW, number of different words.

Table 4

MeasureChatGPTCopilotMagic SchoolGemini
MLUw0.930.930.970.9
MLUm0.910.910.930.89
NTW0.940.940.980.85
NDW0.810.740.970.88
SI0.480.70.430.65
Mazes0.940.930.890.93

Standardized β coefficients by measure and LLM chatbot.

MLUw, mean length of utterance in words; MLUm, mean length of utterance in morphemes; SI, subordination index; NTW, number of total words; NDW, number of different words.

Table 5

LLM Chatbot
PredictorsChatGPTCopilotMagicSchoolGemini
Model 1Model 2Model 1Model 2Model 1Model 2Model 1Model 2
β (SE)β (SE)β (SE)β (SE)β (SE)β (SE)β (SE)β (SE)
Intercept0 (0.28)*0 (0.36)*0 (0.32)0 (0.43)0.24 (0.17)0 (0.21)0 (0.29)***0 (0.33)***
Chatbot MLUw0.93 (0.04)***0.9 (0.05)***0.93 (0.05)***0.91 (0.06)***0.97 (0.03)***0.96 (0.03)***0.9 (0.05)***0.87 (0.05)***
Age group−0.02 (0.09)−0.1 (0.08)*−0.01 (0.05)−0.04 (0.09)
Language group−0.06 (0.08)−0.01 (0.08)−0.05 (0.05)*−0.2 (0.08)***
Age group*MLUw0.06 (0.05)−0.02 (0.06)−0.03 (0.03)−0.01 (0.06)
Language group*MLUw−0.02 (0.04)−0.09 (0.05)0.02 (0.03)0.11 (0.04)*
F501.81104.56501.1109.781,558.51316.79372.896.24
R20.860.870.860.880.950.950.820.86
ΔR20.010.020.000.04

Forward regression models with each LLM chatbot-calculated mean length of utterance in words as the predictor compared with the corresponding SALT microstructure measure as the outcome.

*p < 0.05, **p < 0.01, ***p < 0.001. MLUw, mean length of utterance in words.

ICCs varied across microstructure measures and LLM chatbots (Table 3). SI generally exhibited lower ICC values than the other microstructure measures, whereas the remaining measures showed relatively similar levels of agreement across LLM chatbots.

Standardized β coefficients also varied across measures (Table 4). NTW and MLUw demonstrated consistently high association with SALT (βs ≥0.90), whereas SI β coefficients ranged from 0.43 to 0.70. MLUm, NDW, and mazes yielded β coefficients that varied more across LLM chatbots.

4.2 Effects of age group and language group on LLM–SALT associations

Results from “model 2” in Tables 5 through 10 address RQ3. For MLUw, no significant Age Group × MLUw or Language Group × MLUw interactions were observed for ChatGPT, Copilot, or MagicSchool, indicating that the association between LLM chatbot-calculated and SALT-calculated MLUw did not differ by age or language group. For Gemini, the Age Group × MLUw interaction was not significant; however, the Language Group × MLUw interaction was significant (p < 0.05), indicating that the association between Gemini-calculated MLUw and SALT MLUw differed by language group. The association was stronger for bilinguals than for monolinguals (see Table 5).

For MLUm, no significant Age Group × MLUm or Language Group × MLUm interactions were observed for ChatGPT, Copilot, and MagicSchool. For Gemini, the Age Group × MLUm interaction was not significant; however, the Language Group × MLUm interaction was significant (p < 0.01), indicating that the association between Gemini-calculated MLUm and SALT MLUm differed by language group. The association was stronger for bilinguals than for monolinguals (see Table 6).

Table 6

LLM Chatbot
ChatGPTCopilotMagicSchoolGemini
Model 1Model 2Model 1Model 2Model 1Model 2Model 1Model 2
Predictorsβ (SE)β (SE)β (SE)β (SE)β (SE)β (SE)β (SE)β (SE)
Intercept0 (0.33)**0 (0.41)**0 (0.38)0 (0.52)0 (0.28)**0 (0.29)***0 (0.32)***0 (0.38)***
Chatbot MLUm0.91 (0.04)***0.89 (0.05)***0.91 (0.04)***0.88 (0.06)***0.93 (0.04)***0.86 (0.04)***0.89 (0.05)***0.86 (0.05)***
Age group−0.03 (0.1)−0.07 (0.11)−0.12 (0.08)**-−0.03 (0.11)
Language group−0.02 (0.09)−0.03 (0.09)−0.11 (0.07)**-−0.14 (0.09)**
Age group*MLUm0.08 (0.06)0.001 (0.06)0.01 (0.04)-−0.01 (0.06)
Language group*MLUm0.01 (0.04)−0.03 (0.05)0.03 (0.04)-0.13 (0.04)**
F404.5181.66395.7978.19553.79138.03317.3377.02
R20.830.840.830.830.870.90.790.83
ΔR20.010.000.030.04

Forward regression models with each LLM chatbot-calculated mean length of utterance in morphemes as the predictor compared with the corresponding SALT microstructure measure as the outcome.

*p < 0.05, **p < 0.01, ***p < 0.001. MLUm, mean length of utterance in morphemes.

For NTW, no significant Age Group × NTW or Language Group × NTW interactions were observed for ChatGPT, Copilot, and MagicSchool. For Gemini, the Age Group × NTW interaction was not significant; however, the Language Group × NTW interaction was significant (p < 0.05), indicating that the association between Gemini-calculated NTW and SALT NTW varied by language group. The association was stronger for bilinguals than for monolinguals (see Table 7).

Table 7

LLM Chatbot
PredictorsChatGPTCopilotMagicSchoolGemini
Model 1Model 2Model 1Model 2Model 1Model 2Model 1Model 2
β (SE)β (SE)β (SE)β (SE)β (SE)β (SE)β (SE)β (SE)
Intercept0 (3.72)0 (3.99)0 (3.77)0 (3.8)*0 (2.48)0 (2.7)0 (5.6)***0 (5.47)**
Chatbot NTW0.94 (0.04)***0.92 (0.04)***0.94 (0.04)***0.9 (0.04)***0.98 (0.03)***0.97 (0.03)***0.85 (0.06)***0.86 (0.07)***
Age group0.01 (1.24)−0.12 (1.15)**0.01 (0.79)0.05 (1.84)
Language group−0.1 (1.23)*−0.06 (1.23)−0.04 (0.81)−0.26 (1.68)***
Age group*NTW−0.03 (0.04)−0.06 (0.04)−0.04 (0.03)−0.02 (0.07)
Language group*NTW−0.03 (0.04)−0.07 (0.04)−0.01 (0.03)0.13 (0.06)*
F665.01139.44641.94151.231,751.6354.44208.0962.12
R20.890.90.890.910.960.960.720.8
ΔR20.010.020.000.08

Forward regression models with each LLM chatbot-calculated number of total words as the predictor compared with the corresponding SALT microstructure measure as the outcome.

*p < 0.05, **p < 0.01, ***p < 0.001. NTW, number of total words.

For NDW, no significant Age Group × NDW or Language Group × NDW interactions were observed for MagicSchool and Gemini. For ChatGPT, the Age Group × NDW interaction was not significant; however, the Language Group × NDW interaction was significant (p < 0.001), indicating that the association between ChatGPT-calculated NDW and SALT NDW varied by language group. The association was stronger for bilinguals than for monolinguals (see Table 8). For Copilot, the Age Group × NDW and Language Group × NDW interactions were significant (ps < 0.001), indicating that the association between Copilot-calculated NDW and SALT NDW differed by age and language group. The association was stronger for bilinguals than for monolinguals, and the association was stronger for school-age children than for preschool children (see Table 8).

Table 8

LLM chatbot
PredictorsChatGPTCopilotMagicSchoolGemini
Model 1Model 2Model 1Model 2Model 1Model 2Model 1Model 2
β (SE)β (SE)β (SE)β (SE)β (SE)β (SE)β (SE)β (SE)
Intercept0 (2.41)***0 (2.61)***0 (2.89)***0 (3.05)**0 (1.32)0 (1.52)0 (2.24)***0 (2.6)***
Chatbot NDW0.81 (0.05)***0.88 (0.06)***0.74 (0.06)***1 (0.07)***0.97 (0.03)***0.96 (0.03)***0.88 (0.05)***0.84 (0.07)***
Age group0.03 (0.89)−0.05 (0.83)−0.003 (0.41)−0.03 (0.76)
Language group−0.07 (0.86)0.07 (0.93)−0.02 (0.43)−0.05 (0.79)
Age group*NDW0.04 (0.05)−0.24 (0.06)***−0.03 (0.03)−0.08 (0.06)
Language group*NDW0.24 (0.05)***0.39 (0.07)***0.01 (0.03)0.04 (0.06)
F161.9842.43100.6943.571,199.18232.64280.656.12
R20.660.730.550.740.940.940.770.78
ΔR20.070.190.00.01

Forward regression models with each LLM chatbot-calculated number of different words as the predictor compared with the corresponding SALT microstructure measure as the outcome.

*p < 0.05, **p < 0.01, ***p < 0.001. NDW, number of different words.

For SI, no significant Age Group × SI or Language Group × SI interactions were observed for ChatGPT, Copilot, and Gemini. For MagicSchool, the Age Group × SI interaction was not significant; however, the Language Group × SI interaction was significant (p < 0.01), indicating that the association between MagicSchool-calculated SI and SALT SI differed by language group. The association was stronger for bilinguals than for monolinguals (see Table 9).

Table 9

LLM CHAtbot
PredictorsChatGPTCopilotMagicSchoolGemini
Model 1Model 2Model 1Model 2Model 1Model 2Model 1Model 2
β (SE)β (SE)β (SE)β (SE)β (SE)β (SE)β (SE)β (SE)
Intercept0 (0.08)***0 (0.08)***0 (0.1)*0 (0.13)*0 (0.12)***0 (0.13)***0 (0.08)***0 (0.08)***
Chatbot SI0.48 (0.07)***0.36 (0.07)***0.7 (0.08)***0.68 (0.11)***0.43 (0.1)***0.47 (0.12)***0.65 (0.07)***0.55 (0.07)***
Age group−0.36 (0.01)***−0.09 (0.01)-−0.3 (0.01)**−0.26 (0.01)**
Language group−0.09 (0.01)−0.03 (0.01)−0.01 (0.02)−0.12 (0.01)
Age Group*SI0.11 (0.07)0.01 (0.11)0.03 (0.1)−0.16 (0.08)
Language Group*SI0.14 (0.07)0.15 (0.09)0.32 (0.11)**0.15 (0.08)
F24.9510.277.0316.5218.639.8360.117.01
R20.230.40.480.510.190.390.420.52
ΔR20.170.030.20.1

Forward regression models with each LLM chatbot-calculated subordination index as the predictor compared with the corresponding SALT microstructure measure as the outcome.

*p < 0.05, **p < 0.01, ***p < 0.001. SI, subordination index.

For mazes, no significant Age Group × mazes or Language Group × mazes interactions were observed for ChatGPT, Copilot, and Gemini. For MagicSchool, the Age Group × mazes interaction was not significant; however, the Language Group × mazes interaction was significant (p < 0.05), indicating that the association between MagicSchool-calculated mazes and SALT mazes varied by language group. The association was stronger for monolinguals than for bilinguals (see Table 10).

Table 10

LLM Chatbot
PredictorsChatGPTCopilotMagicSchoolGemini
Model 1Model 2Model 1Model 2Model 1Model 2Model 1Model 2
β (SE)β (SE)β (SE)β (SE)β (SE)β (SE)β (SE)β (SE)
Intercept0 (0.25)0 (0.28)0 (0.28)0 (0.31)0 (0.34)0 (0.34)0 (0.28)0 (0.31)*
Chatbot Mazes0.94 (0.04)***0.96 (0.06)***0.93 (0.04)***0.93 (0.06)***0.89 (0.05)***0.8 (0.06)***0.93 (0.04)***0.98 (0.06)***
Age group0.04 (0.16)0.05 (0.17)0.03 (0.2)−0.06 (0.17)
Language group0.04 (0.18)0.04 (0.19)−0.11 (0.21)*0.12 (0.21)*
Age Group*Mazes0.02 (0.04)0.04 (0.05)0.07 (0.05)−0.03 (0.04)
Language group*Mazes0.001 (0.06)−0.03 (0.06)−0.13 (0.06)*−0.01 (0.06)
F657.27129.04534.12108.68306.7871.2517.81112.53
R20.890.890.870.870.790.820.860.88
ΔR20.00.00.030.02

Forward regression models with each LLM chatbot-calculated mazes as the predictor compared with the corresponding SALT microstructure measure as the outcome.

*p < 0.05, **p < 0.01, ***p < 0.001.

4.3 Effect of prompt specificity on MLUm and SI associations

Multiple regression models for RQ4 are included in Table 11. Each model included the LLM chatbot-calculated microstructure measure, prompt number, and their interaction as predictors of the corresponding SALT-calculated microstructure measure. Interpretation centers on the interaction term, which tests whether prompt specificity alters the magnitude of this association.

Table 11

LLM chatbot
PredictorsChatGPTCopilotMagicSchoolGemini
β (SE)β (SE)β (SE)β (SE)
Intercept0 (0.2)***0 (0.26)***0 (0.28)***0 (0.24)***
Chatbot MLUm0.94 (0.03)***0.91 (0.03)***0.88 (0.04)***0.89 (0.03)***
Prompt #0.04 (0.05)−0.12 (0.06)***0.15 (0.08)***−0.07 (0.07)
MLUm*prompt #−0.01 (0.03)0.1 (0.03)**0.18 (0.04)***−0.02 (0.03)
F395.9217.33141.54200.24
R20.880.80.720.79
Predictorsβ (SE)β (SE)β (SE)β (SE)
Intercept0 (0.05)***0 (0.04)***0 (0.03)***0.7 (0.03)***
Chatbot SI0.62 (0.04)***0.85 (0.004)***0.57 (0.004)***−0.1 (0.005)***
Prompt #−0.26 (0.01)***−0.44 (0.01)***−0.31 (0.01)***−0.11 (0.01)
SI*Prompt #−0.05 (0.04)0.14 (0.004)***−0.04 (0.004)0.1 (0.005)
F26.7348.416.542.57
R20.330.470.230.44

Multiple regression models comparing prompt 1 & prompt 2 with MLUm and SI entered individually for each LLM chatbot compared with the corresponding SALT microstructure measure.

*p < 0.05, **p < 0.01, ***p < 0.001. MLUm, mean length of utterance in morphemes; SI, subordination index.

For MLUm, no significant prompt × MLUm interactions were observed for ChatGPT or Gemini. Significant interactions were observed for Copilot (p < 0.01) and MagicSchool (p < 0.001), indicating that prompt specificity altered the association between the LLM chatbot-calculated MLUm and the corresponding SALT-calculated MLUm.

For SI, no significant prompt × SI interactions were observed for ChatGPT, MagicSchool, or Gemini. A significant prompt × SI interaction was observed for Copilot (p < 0.001), indicating that prompt specificity altered the association between the LLM chatbot-calculated SI and the corresponding SALT-calculated SI.

5 Discussion

This exploratory study examined the association and agreement between microstructure measures calculated by four LLM chatbots (ChatGPT, Copilot, MagicSchool, and Gemini) and those calculated using SALT. English language samples were collected from 84 children (35 girls; M = 75.62 months, SD = 22.51), including 30 Japanese-English bilinguals and 54 English monolinguals, ages 2–9 years. SALT was used to calculate MLUw, MLUm, SI, NTW, NDW, and number of mazes. Each transcript was then broken into C-units and analyzed by the four LLM chatbots using an initial prompt, followed by a second, more specific prompt that included MLUm and SI calculation parameters to examine the effect of prompt specificity. Key findings were: (1) all LLM chatbot-calculated microstructure measures were significantly associated with the corresponding SALT measures; (2) NTW and MLUw generally showed higher associations with SALT, whereas SI exhibited lower associations. Agreement, as measured by ICCs, varied across LLM chatbots, with mazes showing the least variability; (3) age group significantly moderated the association only for Copilot-calculated NDW, with a stronger association with SALT NDW for school-age children than for preschool children; (4) language group significantly moderated the associations for ChatGPT & Copilot NDW, Magic school SI & mazes, and Gemini MLUw, MLUm, & NTW, with the associations being stronger for the bilingual group for all except MagicSchool mazes; and (5) greater prompt specificity did not improve associations between LLM chatbot-calculated and SALT-calculated measures, with weaker associations observed for Copilot MLUm and SI and MagicSchool MLUm.

5.1 Association & agreement between LLM and SALT microstructure measures

Across LLM chatbots, all LLM chatbot-calculated microstructure measures were significantly associated with the corresponding SALT measure. MLUw and NTW generally showed the strongest and most consistent associations, while MLUm, NDW, and mazes were also strongly associated, although the magnitude of the associations varied across LLM chatbots. SI showed lower associations. Agreement likewise varied across microstructure measures and LLM chatbots, indicating that the degree of agreement with SALT depended on both the measure and the LLM chatbot. Consistent with the association findings, MLUw and NTW generally exhibited higher agreement, whereas SI exhibited lower agreement. MLUm, NDW, and mazes showed greater variability in agreement across LLM chatbots. Overall, association and agreement showed broadly similar descriptive patterns across measures, although the magnitude of the coefficients varied.

These findings support RQ1, indicating that LLM chatbot-calculated microstructure measures are associated and agree with SALT microstructure measures, although the strength of association and agreement depends on the microstructure measure. Notably, these results suggest improvements in the calculation capabilities of LLM chatbots compared with the findings of , who reported smaller ICCs and Pearson correlation coefficients between ChatGPT and SALT for MLUw and NDW than the current study. This improvement likely reflects advances in ChatGPT, as their study used ChatGPT 4o in 2024, whereas the present study used ChatGPT 5 in late 2025. Additionally, we separated utterances into C-units, which likely assisted the LLM chatbots when calculating MLUw. Overall, the present findings demonstrate convergence under a highly controlled analytic workflow. They do not establish that LLM chatbot-calculated scores are interchangeable with SALT, nor do they demonstrate clinical decision accuracy.

For RQ2, consistent with our hypothesis, SI yielded lower association and agreement with SALT than the other microstructure measures. Contrary to our expectation, however, MLUm also demonstrated relatively high association and agreement with SALT, comparable to measures requiring relatively simple word or morpheme counts. One possible explanation is that SI requires identifying clause boundaries and distinguishing main and subordinate clauses, which involves more complex linguistic analysis than measures requiring relatively simple word or morpheme counts (e.g., SALT scoring procedures; Miller et al., 2019). However, it is important to note that large standardized β coefficients indicate close linear alignment within this sample but do not establish measurement equivalence or clinical interchangeability. Agreement also varied across LLM chatbots and microstructure measures, indicating that the degree of agreement with SALT depended on both the measure and the chatbot. These findings are consistent with research showing that numerical or rule-based computations produced by LLM chatbots can vary in reliability (e.g., ; ; King, 2023). Additionally, our findings of improved MLUw calculations compared to , as stated earlier, could be due to how we prepared the transcripts. We broke up the transcripts in the current study by C-units, which may have supported the calculation accuracy of MLUw. It is possible that, if we had provided more information about the number of main and subordinate clauses in the transcripts themselves, rather than as context within the second prompt, the calculation accuracy would have gone up. This brings up an interesting question regarding how to prompt LLM chatbots for LSA: Is it more effective to add context to the prompt, or is it more effective to add notes and codes within the transcript and stick to a simpler, more general prompt? In either case, adding too much information into the prompt or transcript itself, such as the types of clauses or separating morphemes, would negate the purpose of trying to save time by using LLM chatbots to code and calculate microstructure measures. Thus, it is possible that LLMs currently do not have the capability to accurately and consistently discern and code different clauses.

Interestingly, descriptive differences in performance were observed across the LLM chatbots. For example, MagicSchool generally showed higher association and agreement with SALT for several measures, such as NTW (β = 0.98 and ICC(2,1) = 0.97), MLUw (β = 0.97 and ICC(2,1) = 0.97), NDW (β = 0.97 and ICC(2,1) = 0.97), and MLUm (β = 0.93 and ICC(2,1) = 0.93), while ChatGPT exhibited the highest association for mazes (β = 0.94 and ICC(2,1) = 0.93). This variation may reflect differences in model parameters or training data. However, differences in model performance also touches on the “black box” problem. Specifically, it is difficult to explain why each model is performing differently, as it could be the training data, parameters that control how the model learns from the data, or something entirely different. The opaqueness of these models makes it difficult to explain differences in performance or even choose a model that best suits a specific task. Regardless, we did not use formal statistical tests to compare calculation performance across LLM chatbots. The purpose of the current study was to evaluate the association and agreement between each LLM chatbot and SALT, rather than to determine whether one chatbot outperformed another. Consequently, the observed differences across chatbots should be interpreted descriptively and warrant confirmation in future studies designed for direct model comparisons.

The hypothesis regarding age group was largely not supported, whereas language group moderated the associations for several microstructure measures. For age groups, no evidence of moderation was observed across chatbots, with the exception of Copilot NDW. Specifically, Copilot NDW showed a stronger association for school-age children than for preschool children. This difference may reflect distributional differences in NDW between age groups, including potential differences in variance or score range, which can influence standardized slope estimates. However, aside from Copilot-calculated NDW, age-related differences in coefficient magnitude were minimal. To our knowledge, no research has examined age as a moderating variable in AI-assisted narrative analysis. Jones et al. (2019) evaluated ML for automating narrative macrostructural analysis for children ages five to nine, and they found that BERT could reliably score narrative macrostructure elements at levels comparable to human raters, while examined ChatGPT-calculated microstructure measures for narratives of children ages six to eight, although neither of these studies included age as a moderating variable. Future research should examine age as a moderating variable in AI-assisted LSA.

For language groups, all the LLM chatbots showed a significant difference between the strength of association with SALT for at least one microstructure measure. For ChatGPT and Copilot, NDW yielded greater associations for the bilingual group. For Magic school, the SI score yielded a larger coefficient for bilingual children, whereas mazes yielded a larger coefficient for monolingual children. Finally, for Gemini, the MLUw, MLUm, and NTW scores yielded larger coefficients for the bilingual group. Overall a notable pattern was that the majority of the LLM chatbots had stronger association for the bilingual group, with MagicSchool maze calculations being the exception. These findings do not indicate differential computational accuracy across groups. One possible explanation for the observed pattern may be related to the sample sizes (30 bilinguals vs. 54 monolinguals) which may have introduced greater variation in the monolingual groups' data. It also may be explained by the length of the groups' samples. The bilinguals produced shorter language samples (NTW M = 84.33) compared to the monolinguals (NTW M = 104.72), which also may have introduced greater variation and contributed to the observed differences in the strength of the associations. Additionally, the data in the current study was used from , who found subtle differences between the English narratives of the bilingual and monolingual groups, which also could account for language group modulating LLM chatbot performance. However, the current study did find generally that LLM chatbots are not significantly less associated with SALT when handling the calculations of bilinguals in English, which is a positive, albeit preliminary finding. These findings reflect differences in coefficient magnitude within this sample and do not indicate differential computational accuracy across groups. Thus, future studies should further examine bilingualism as a moderating variable in LLM chatbot LSA.

Finally, our hypothesis for RQ4 was not supported. None of the LLM chatbot-calculated MLUm or SI scores significantly increased in strength of association with SALT scores when prompted with more specific calculation parameters. In opposition to our hypothesis, for Copilot and MagicSchool MLUm and Copilot SI, the standardized coefficients were smaller under the second prompt. This finding was surprising given that research on “prompt engineering” or “adding context” has shown that adding details to the prompt can improve the output (; Kirilenko, 2026; Oppenlaender et al., 2025). However, Oppenlaender et al. (2025) also found that prompt engineering is not necessarily an innate process but instead requires individual experimentation. Additionally, several prompts may be required to receive the most accurate response (). Another reason for the differences in performance across prompts and LLM chatbots may have to do with their respective tokenizers. A tokenizer is a tool that breaks text into smaller units called “tokens,” such as words or characters, before converting them into numbers that an LLM chatbot can process. LLMs have context window limitations, so users can have conversation with a limited number of words. When that limit is surpassed, the model starts to forget about context, and exceeding the tokenizer limit can lead to “hallucinations,” or inaccuracies in the LLM chatbot's output (; Hicks et al., 2024). The longer prompt may have increased the workload of the LLM chatbots' tokenizer, increasing the potential for errors (; Kim et al., 2025). Future studies should explore different prompts to improve the accuracy of LLM chatbot SI calculations.

Overall, the current study identified substantial correspondence between LLM chatbot-calculated and SALT microstructure measures under a controlled analytic protocol, particularly for MLUw (βs = 0.90–0.97) and NTW (βs = 0.92–0.98), MLUm (βs = 0.89–0.93), NDW (βs = 0.74–0.97), and mazes (βs = 0.89–0.94). On the other hand, SI (βs = 0.43–0.70) yielded comparatively weaker associations. Agreement for each microstructure measure was much more variable, indicating that high association with SALT did not necessarily correspond to equally high levels of agreement. These findings suggest that some measures may have been calculated consistently yet still differed systematically from the corresponding SALT values. Although association and agreement generally followed similar trends across measures, they provide complementary information and should not be interpreted as interchangeable indicators of correspondence. The coefficients in the current study appeared stronger for some of the measures than those from , suggesting improvement in LLM chatbot calculation accuracy. Additionally, Organisciak and Acar (2026) point out that LLM chatbots are designed to deal with text. Specifically, they deal in the currency of discrete chunks of text that are predicted in sequence. They are not designed to deal with quantitative analysis, since they operate by relying on statistical patterns and probabilities to generate text, rather than verifying truthfulness (). Thus, it is essential that more research is done with LLM chatbots to improve quantitative output, similar to the research done with LLM chatbot-assisted qualitative text analysis (Jenner et al., 2025; Tai et al., 2024), to uncover a reliable method for using LLM chatbots to assist with LSA.

Notwithstanding the strong agreement and association found in the current study, because relatively small differences in microstructure scores may influence clinical interpretation, these findings should not be construed as evidence of diagnostic equivalence. Within this sample, certain measures yielded comparatively larger coefficients, but these descriptive differences require replication and further validation.

5.2 Limitations

Several limitations should be considered. First, the sample consisted of 84 transcripts from typically developing children ages two to nine. The number of participants could be increased, and the age range could also be broadened to draw stronger conclusions. Additionally, our study did not include children with a language impairment, which limits the generalizability of the current study's findings to clinical populations and how LLM chatbots can assist with clinical decision making. Furthermore, the recommended number of utterances for a language sample is around 50 utterances, or a 3-min sample (). The average number of utterances in the current study was only 14.88. We noted in the discussion that the LLM chatbots were more accurate at calculating several microstructure measures for the bilingual group, possibly because they produced shorter language samples. The shorter samples may have reduced variation and reduced the opportunity for calculation errors. Thus, it is important that future research examines how transcript length impacts the accuracy of LLM chatbot-assisted LSA persists with substantially longer samples. Another limitation to this study is that we did not thoroughly explore the reasons why each LLM chatbot was not as accurate as SALT. Future research should probe the LLM chatbots with why the microstructure calculations turned out as they did, then compare the LLM chatbot's output to the SALT output to tease apart the specific areas of accuracy/inaccuracy. It would also be useful to assess if the SALT measures and corresponding LLM measures lead to different diagnostic decisions due to minor differences in the output. Additionally, we compared LLM chatbot-calculated microstructure measures to the SALT-calculated equivalents completed by the PI and several RAs, but future research should compare LLM chatbot performance with SLP analysis of the same microstructure measures. Finally, a major limitation to our study may be replicability. While future studies can use the same methods to prompt each LLM chatbot, the same performance from each LLM chatbot cannot be guaranteed due to continual changing and updating. This limitation demonstrates the “blackbox” problem—individuals using LLM chatbots, who are not involved in their development, may not know the complex innerworkings of the LLM chatbots. The data was collected for this study in September, October, and November of 2025, but these LLM chatbots may have different model specifications now and in the future. This problem also applies to clinical work. LLM models may lack stability over time to reliably incorporate them into assessment and therapy, as constant change would bias the output. Therefore, this study only serves as a starting point for future research examining the clinical utility of LLM chatbots.

5.3 Clinical implications

This study has several implications, the most important being that this was an exploratory study, and these findings should be interpreted as descriptive rather than evidence of LLM chatbot superiority or workflow efficiency. Our findings suggest that LLM chatbots may have greater potential to assist with the calculation of measures requiring relatively simple word or morpheme counts (e.g., MLUw and MLUm, as well as NTW and NDW) than more linguistically complex measures such as SI. This distinction is clinically important because it identifies where LLM-assisted LSA may be most feasible while reinforcing that more complex syntactic measures still require additional validation. Future research should compare chatbot performance, improve calculation accuracy, and evaluate the practical benefits of LLM-assisted LSA to thoroughly improve implementation of LLM chatbots into clinical practice.

Implementing LLM chatbots into SLP practices can help with a variety of tasks (; ; ; ; Lewis et al., 2025). However, SLPs may rely too much on LLM chatbots (), thus, it is important that researchers thoroughly examine the validity and efficacy of LLM chatbots for research, education, and clinical settings, as there is a lack of externally validated studies on the implementation of AI into education and healthcare (World Health Organization, 2021; Villamil et al., 2020). Higher education institutions should become responsible for developing critical thinking, ethical literacy, and a healthy skepticism toward LLMs (), and SLPs and educators should continue to educate themselves on the limitations of LLM chatbots, such as by using guided rubrics for evaluating AI tools for SLP practice (). LLM chatbots still have serious limitations (e.g., ; Hicks et al., 2024), so they should be implemented cautiously, and the output should always be double checked ().

Statements

Data availability statement

The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.

Ethics statement

The studies involving humans were approved by University of Colorado institutional review board. The studies were conducted in accordance with the local legislation and institutional requirements. Written informed consent for participation in this study was provided by the participants' legal guardians/next of kin.

Author contributions

RH: Data curation, Writing – review & editing, Conceptualization, Writing – original draft, Investigation, Project administration, Formal analysis. PK: Conceptualization, Writing – review & editing, Supervision, Investigation, Writing – original draft, Formal analysis.

Funding

The author(s) declared that financial support was not received for this work and/or its publication.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

The author PK declared that they were an editorial board member of Frontiers, at the time of submission. This had no impact on the peer review process and the final decision.

Generative AI statement

The author(s) declared that Generative AI was used in the creation of this manuscript. Generative AI was used to reduce the word count.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/flang.2026.1929126/full#supplementary-material

References

  • 1

    AkogluH. (2018). User's guide to correlation coefficients. Turk. J. Emerg. Med.18, 9193. doi: 10.1016/j.tjem.2018.08.001

  • 2

    AllenA. K.BrennanC.RisemanC.KleiberH.HilgerA. I. (2025). Harnessing AI for aphasia: a case report on ChatGPT's role in supporting written expression. Front. Rehabil. Sci.6:1600145. doi: 10.3389/fresc.2025.1600145

  • 3

    American Speech-Language-Hearing Association (2004). Preferred Practice Patterns for the Profession of Speech-Language Pathology. Rockville, MD: American Speech-Language-Hearing Association.

  • 4

    AsciF.MarsiliL.SuppaA.SaggioG.MichettiE.Di LeoP.et al. (2023). Acoustic analysis in stuttering: a machine-learning study. Front. Neurol.14:1169707. doi: 10.3389/fneur.2023.1169707

  • 5

    AustinJ.BenasK.CaicedoS.ImiolekE.PiekutowskiA.GhanimI. (2025). Perceptions of artificial intelligence and ChatGPT by speech-language pathologists and students. Am. J. Speech Lang. Pathol.34, 174200. doi: 10.1044/2024_AJSLP-24-00218

  • 6

    AzariaA. (2022). ChatGPT Usage and Limitations. Available online at: https://hal.science/hal-03913837 (Accessed November 19, 2025).

  • 7

    BallochJ.SridharanS.OldhamG.WrayJ.GoughP.RobinsonR.et al. (2024). Use of an ambient artificial intelligence tool to improve quality of clinical documentation. Future Healthc J.11:100157. doi: 10.1016/j.fhj.2024.100157

  • 8

    BeccaluvaE. A.CataniaF.ArosioF.GarzottoF. (2024). Predicting developmental language disorders using artificial intelligence and a speech data analysis tool. Hum.-Comput. Interact.39, 842. doi: 10.1080/07370024.2023.2242837

  • 9

    BenwayN. R.PrestonJ. L. (2024). Artificial intelligence–assisted speech therapy for /ɹ/: a single-case experimental study. Am. J. Speech Lang. Pathol.33, 24612486. doi: 10.1044/2024_AJSLP-23-00448

  • 10

    BenwayN. R.PrestonJ. L. (2025). Equipping speech-language clinicians for the critical appraisal of an artificial intelligence–driven, evidence-based future. Lang. Speech Hear. Serv. Sch.56, 442468. doi: 10.1044/2025_LSHSS-24-00085

  • 11

    BhardwajA.SharmaM.KumarS.SharmaS.SharmaP. C. (2024). Transforming pediatric speech and language disorder diagnosis and therapy: the evolving role of artificial intelligence. Health Sci. Rev.12:100188. doi: 10.1016/j.hsr.2024.100188

  • 12

    BirolN. Y.ÇiftciH. B.YilmazA.ÇaglayanA.AlkanF. (2025). Is there any room for ChatGPT AI bot in speech-language pathology?Eur. Arch. Otorhinolaryngol.282, 32673280. doi: 10.1007/s00405-025-09295-y

  • 13

    BottingN. (2002). Narrative as a tool for the assessment of linguistic and pragmatic impairments. Child Lang. Teach. Ther.18, 121. doi: 10.1191/0265659002ct224oa

  • 14

    BoucherP. (2020). Artificial Intelligence: how does it Work, Why Does it Matter, and What We Can Do About it? Brussels: Scientific Foresight Unit (STOA) EPRS | European Parliamentary Research Service.

  • 15

    BrigantiG. (2024). How ChatGPT works: a mini review. Eur. Arch. Otorhinolaryngol.281, 15651569. doi: 10.1007/s00405-023-08337-7

  • 16

    CannizzaroM.WoolpertD. (2024). What ChatGPT Can(‘t) Do When Automatically Analyzing a Language Sample. Boston, MA: ASHA.

  • 17

    CordellaC.MarteM. J.LiuH.KiranS. (2025). An introduction to machine learning for speech-language pathologists: concepts, terminology, and emerging applications. Perspect. ASHA Spec. Interest Groups10, 432450. doi: 10.1044/2024_PERSP-24-00037

  • 18

    DaleP. S. (1991). The validity of a parent report measure of vocabulary and syntax at 24 months. J. Speech Hear. Res.34, 565571. doi: 10.1044/jshr.3403.565

  • 19

    DuY.Juefei-XuF. (2023). Generative AI for therapy? Opportunities and Barriers for ChatGPT in Speech-Language Therapy., (Kigali, Rwanda). Available online at: https://openreview.net/forum?id=cRZSr6Tpr1S (Accessed November 25, 2025).

  • 20

    DuránP.MalvernD.RichardsB.ChipereN. (2004). Developmental trends in lexical diversity. Appl. Linguist.25, 220242. doi: 10.1093/applin/25.2.220

  • 21

    EbertK. D.PhamG. (2017). Synthesizing information from language samples and standardized tests in school-age bilingual assessment. Lang. Speech Hear. Serv. Sch.48, 4255. doi: 10.1044/2016_LSHSS-16-0007

  • 22

    FischerS. (2025). ChatGPT is still by far the most popular AI chatbot. Axios. Available online at: https://www.axios.com/2025/09/06/ai-chatbot-popularity (Accessed December 19, 2025).

  • 23

    FriederS.PinchettiL.ChevalierA.GriffithsR.-R.SalvatoriT.LukasiewiczT.et al. (2023). Mathematical capabilities of ChatGPT. doi: 10.52202/075280-1205

  • 24

    GagarinaN.BohnackerU.LindgrenJ. (2019a). Macrostructural organization of adults' oral narrative texts. ZAS Pap. Linguist.62, 190208. doi: 10.21248/zaspil.62.2019.449

  • 25

    GagarinaN. V.KlopD.KunnariS.TanteleK.VälimaaT.BalčiunieneI.et al. (2019b). MAIN: multilingual assessment instrument for narratives. Zaspil56:155. doi: 10.21248/zaspil.56.2019.414

  • 26

    GeorgiouG. P. (2025). Transforming speech-language pathology with AI: opportunities, challenges, and ethical guidelines. Healthcare13:2460. doi: 10.3390/healthcare13192460

  • 27

    GilardiF.AlizadehM.KubliM. (2023). ChatGPT outperforms crowd workers for text-annotation tasks. Proc. Nat. Acad. Sci.120:e2305016120. doi: 10.1073/pnas.2305016120

  • 28

    GreenJ. R. (2024). Artificial intelligence in communication sciences and disorders: introduction to the forum. J. Speech Lang. Hear. Res.67, 41574161. doi: 10.1044/2024_JSLHR-24-00594

  • 29

    HannaM.YanaS. (2025). Stench of errors or the shine of potential: the challenge of (Ir)responsible use of ChatGPT in speech-language pathology. Int. J. Lang. Commun. Disord60:e70088. doi: 10.1111/1460-6984.70088

  • 30

    HayesR. L.KanP. F. (2026). Shared and divergent patterns in narrative skills: comparing English monolingual and Japanese–English bilingual children. Front. Psychol.17:1747702. doi: 10.3389/fpsyg.2026.1747702

  • 31

    HeilmannJ.MillerJ.NockertsA. (2010a). Sensitivity of narrative organization measures using narrative retells produced by young school-age children. Lang. Test.27, 603626. doi: 10.1177/0265532209355669

  • 32

    HeilmannJ.NockertsA.MillerJ. F. (2010b). Language sampling: does the length of the transcript matter?Lang. Speech Hear. Serv. Sch.41, 393404. doi: 10.1044/0161-1461(2009/09-0023)

  • 33

    HeilmannJ.TucciA.PlanteE.MillerJ. F. (2020). Assessing functional language in school-aged children using language sample analysis. Perspect. ASHA Spec. Interest Groups5, 622636. doi: 10.1044/2020_PERSP-19-00079

  • 34

    HicksM. T.HumphriesJ.SlaterJ. (2024). ChatGPT is bullshit. Ethics Inf. Technol.26:38. doi: 10.1007/s10676-024-09775-5

  • 35

    JennerS.RaidosD.AndersonE.FleetwoodS.AinsworthB.FoxK.et al. (2025). Using large language models for narrative analysis: a novel application of generative AI. Methods Psychol.12:100183. doi: 10.1016/j.metip.2025.100183

  • 36

    JonesS.FoxC.GillamS.GillamR. (2019). An exploration of automated narrative analysis via machine learning. PLoS ONE14:e0224634. doi: 10.1371/journal.pone.0224634

  • 37

    JusticeL. M.BowlesR. P.KaderavekJ. N.UkrainetzT. A.EisenbergS. L.GillamR. B. (2006). The index of narrative microstructure: a clinical tool for analyzing school-age children's narrative performances. Am. J. Speech Lang. Pathol.15, 177191. doi: 10.1044/1058-0360(2006/017)

  • 38

    KapantzoglouM.FergadiotisG.RestrepoM. A. (2017). Language sample analysis and elicitation technique effects in bilingual children with and without language impairment. J. Speech Lang. Hear. Res.60, 28522864. doi: 10.1044/2017_JSLHR-L-16-0335

  • 39

    KempK.KleeT. (1997). Clinical language sampling practices: results of a survey of speech-language pathologists in the United States. Child Lang. Teach. Ther.13, 161176. doi: 10.1177/026565909701300204

  • 40

    KimY.RussellJ.KarpinskaM.IyyerM. (2025). One Ruler to Measure them all: Benchmarking Multilingual Long-Context Language Models. Montreal: COLM.

  • 41

    KingM. (2023). Administration of the text-based portions of a general IQ test to five different large language models. Available online at: https://www.authorea.com/doi/full/10.36227/techrxiv.22645561?commit=34a0376947a0e3d92c87bc3e7f6b546e8d354df3 (Accessed November 11, 2025).

  • 42

    KirilenkoA. P. (2026). “Text analysis with large language models (LLMs),” in Practical Data Mining with AI for Social Scientists, ed. A. P. Kirilenko (Cham: Springer Nature Switzerland), 389431. doi: 10.1007/978-3-031-89689-7_13

  • 43

    KlatteI. S.van HeugtenV.ZwitserloodR.GerritsE. (2022). Language sample analysis in clinical practice: speech-language pathologists' barriers, facilitators, and needs. Lang. Speech Hear. Serv. Sch.53, 116. doi: 10.1044/2021_LSHSS-21-00026

  • 44

    KleeT.StokesS. F.WongA. M.-Y.FletcherP.GavinW. J. (2004). Utterance length and lexical diversity in cantonese-speaking children with and without specific language impairment. J. Speech Lang. Hear. Res.47, 13961410. doi: 10.1044/1092-4388(2004/104)

  • 45

    KohnkeL.MoorhouseB. L.ZouD. (2023). ChatGPT for language teaching and learning. RELC J.54, 537550. doi: 10.1177/00336882231162868

  • 46

    KošutarS.KramarićM.HrŽicaG. (2022). The relationship between narrative microstructure and macrostructure: differences between six- and eight-year-olds. Psychol. Lang. Commun.26, 126153. doi: 10.2478/plc-2022-0007

  • 47

    LammertJ. M.RobertsA. C.McRaeK.BatterinkL. J.ButlerB. E. (2025). Early identification of language disorders using natural language processing and machine learning: challenges and emerging approaches. J. Speech, Lang.Hear. Res.68, 705718. doi: 10.1044/2024_JSLHR-24-00515

  • 48

    LeadholmB. J.MillerJ. F. (1994). Language sample analysis: the wisconsin guide. publication sales, wisconsin department of public instruction, Drawer 179, Milwaukee, WI 53293-0179. Available online at: https://eric.ed.gov/?id=ED371528 (Accessed November 27, 2025).

  • 49

    LeeS.-B.YoonJ.-H. (2025). Current applications and perceptions of generative ai in speech-language pathology clinical practice and research. jslhd34, 155169. doi: 10.15724/jslhd.2025.34.1.155

  • 50

    LewisA.DangolA.SuhH.OlszewskiA.FogartyJ.KientzJ. A. (2025). Exploring AI-Based Support in Speech-Language Pathology for Culturally and Linguistically Diverse Children., in proceedings of the 2025 chi conference on human factors in computing systems, (New York, NY, USA: Association for Computing Machinery), 119. doi: 10.1145/3706598.3714131

  • 51

    LiuH.MacWhinneyB.FrommD.LanziA. (2023). Automation of language sample analysis. J. Speech, Lang.Hear. Res.66, 24212433. doi: 10.1044/2023_JSLHR-22-00642

  • 52

    LüdtkeU.BornmanJ.De WetF.HeidU.OstermannJ.RumbergL.et al. (2023). Multidisciplinary perspectives on automatic analysis of children's language samples: where do we go from here?Folia Phoniatr. Logop.75, 112. doi: 10.1159/000527427

  • 53

    MacWhinneyB. (2000). The CHILDES project: Tools for analyzing talk: Transcription format and programs,Vol. 1, 3rd Edn. Mahwah, NJ: Lawrence Erlbaum Associates Publishers.

  • 54

    MacWhinneyB.FrommD. (2022). Language sample analysis with talkbank: an update and review. Front. Commun.7:865498. doi: 10.3389/fcomm.2022.865498

  • 55

    MarchmanV. A.Martínez-SussmannC. (2002). Concurrent validity of caregiver/parent report measures of language for children who are learning both English and Spanish. J. Speech Lang. Hear. Res.45, 983997. doi: 10.1044/1092-4388(2002/080)

  • 56

    MillerJ. F.AndriacchiK.NockertsA. (eds.) (2019). Assessing Language Production Using SALT, 3rd Edn. Madison, WI: SALT Software LLC.

  • 57

    MillerJ. F.ChapmanR. S. (1981). The relation between age and mean length of utterance in morphemes. J. Speech Lang. Hear. Res.24, 154161. doi: 10.1044/jshr.2402.154

  • 58

    MislevyJ. L.RuppA. A. (2010). “Concurrent Validity,” in Encyclopedia of Research Design, ed. N. Salkind (Thousand Oaks, CA: SAGE Publications, Inc).

  • 59

    OppenlaenderJ.LinderR.SilvennoinenJ. (2025). Prompting AI art: an investigation into the creative skill of prompt engineering. Int. J. Hum.–Comput. Interact. 41, 1020710229. doi: 10.1080/10447318.2024.2431761

  • 60

    OrganisciakP.AcarS. (2026). Know when to trust: making AI scoring more reliable for educational assessment. Behav. Res.58:178. doi: 10.3758/s13428-026-03058-1

  • 61

    OtwinowskaA.MieszkowskaK.Białecka-PikulM.OpackiM.HamanE. (2020). Retelling a model story improves the narratives of polish-english bilingual children. Int. J. Biling. Educ. Biling.23, 10831107. doi: 10.1080/13670050.2018.1434124

  • 62

    OwenA. J.LeonardL. B. (2002). Lexical diversity in the spontaneous speech of children with specific language impairment. J. Speech, Lang. Hear. Res.45, 927937. doi: 10.1044/1092-4388(2002/075)

  • 63

    ParkerM. D.BrorsonK. (2005). A comparative study between mean length of utterance in morphemes (MLUm) and mean length of utterance in words (MLUw). First Lang.25, 365376. doi: 10.1177/0142723705059114

  • 64

    PavelkoS. L.OwensR. E.IrelandM.Hahs-VaughnD. L. (2016). Use of language sample analysis by school-based SLPs: results of a nationwide survey. Lang. Speech Hear. Serv. Sch.47, 246258. doi: 10.1044/2016_LSHSS-15-0044

  • 65

    PokrivcakovaS. (2019). Preparing teachers for the application of AI-powered technologies in foreign language education. J. Lang. Cult. Educ.7, 135153. doi: 10.2478/jolace-2019-0025

  • 66

    RahmanA.RajA.TomyP.HameedM. S. (2024). A comprehensive bibliometric and content analysis of artificial intelligence in language learning: tracing between the years 2017 and 2023. Artif. Intell. Rev.57:107. doi: 10.1007/s10462-023-10643-9

  • 67

    RameshV.AssafR. (2021). Detecting Autism Spectrum Disorders with Machine Learning Models Using Speech Transcripts. Ithaca, NY: Cornell University.

  • 68

    RamosM. N.CollinsP.PeñaE. D. (2022). Sharpening our tools: a systematic review to identify diagnostically accurate language sample measures. J. Speech, Lang. Hear. Res.65, 38903907. doi: 10.1044/2022_JSLHR-22-00121

  • 69

    RiceM. L.SmolikF.PerpichD.ThompsonT.RyttingN.BlossomM. (2010). Mean length of utterance levels in 6-month intervals for children 3 to 9 years with and without language impairments. J. Speech Lang. Hear. Res.53, 333349. doi: 10.1044/1092-4388(2009/08-0183)

  • 70

    ShroutP. E.FleissJ. L. (1979). Intraclass correlations: uses in assessing rater reliability. Psychol. Bull.86, 420428. doi: 10.1037/0033-2909.86.2.420

  • 71

    TaiR. H.BentleyL. R.XiaX.SittJ. M.FankhauserS. C.Chicas-MosierA. M.et al. (2024). An examination of the use of large language models to aid analysis of textual data. Int. J. Qual. Methods23:16094069241231168. doi: 10.1177/16094069241231168

  • 72

    ThalD. J.O'HanlonL.ClemmonsM.FralinL. (1999). Validity of a parent report measure of vocabulary and syntax for preschool children with language impairment. J. Speech Lang. Hear. Res.42, 482496. doi: 10.1044/jslhr.4202.482

  • 73

    VillamilV.DeloriaR.WolbringG. (2020). Artificial intelligence and machine learning: what is the role of social workers, occupational therapists, audiologists, nurses and speech language pathologists according to academic literature and canadian newspaper coverage?, in proceedings of the 5th workshop on icts for improving patients rehabilitation research techniques, (New York, NY: Association for Computing Machinery) 83–87. doi: 10.1145/3364138.3364158

  • 74

    WetherbyA. M.AllenL.ClearyJ.KublinK.GoldsteinH. (2002). Validity and reliability of the communication and symbolic behavior scales developmental profile with very young children. J. Speech Lang. Hear. Res.45, 12021218. doi: 10.1044/1092-4388(2002/097)

  • 75

    WilderA.RedmondS. M. (2024). Updates on clinical language sampling practices: a survey of speech-language pathologists practicing in the United States. Lang. Speech Hear. Serv. Sch.55, 11511166. doi: 10.1044/2024_LSHSS-24-00035

  • 76

    World Health Organization, (2021). Ethics and Governance of Artificial Intelligence for Health: WHO Guidance, 1st Edn. Geneva: World Health Organization.

  • 77

    ZhangM.TangE.DingH.ZhangY. (2024a). Artificial intelligence and the future of communication sciences and disorders: a bibliometric and visualization analysis. J. Speech Lang. Hear. Res.67, 43694390. doi: 10.1044/2024_JSLHR-24-00157

  • 78

    ZhangR.ZouD.ChengG. (2024b). A review of chatbot-assisted learning: pedagogical approaches, implementations, factors leading to effectiveness, theories, and future directions. Interact. Learn. Environ.32, 45294557. doi: 10.1080/10494820.2023.2202704

Summary

Keywords

artificial intelligence, ChatGPT, child language evaluation, language sample analysis, large language models, speech-language pathology

Citation

Hayes RL and Kan PF (2026) Convergence between AI chatbot-calculated microstructure measures and SALT analysis of child language samples. Front. Lang. Sci. 5:1929126. doi: 10.3389/flang.2026.1929126

Received

05 July 2026

Revised

12 August 2026

Accepted

17 August 2026

Published

07 September 2026

Volume

5 - 2026

Edited by

Guillaume Thierry, Bangor University, United Kingdom

Reviewed by

Halima Sahraoui, University of Toulouse, France

Leticia Correa Celeste, University of Brasilia, Brazil

Updates

Copyright

*Correspondence: Richy Lewis Hayes,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics