Abstract
Autism spectrum disorder (ASD) is diagnosed on the basis of speech and communication differences, amongst other symptoms. Since conversations are essential for building connections with others, it is important to understand the exact nature of differences between autistic and non-autistic verbal behaviour and evaluate the potential of these differences for diagnostics. In this study, we recorded dyadic conversations and used automated extraction of speech and interactional turn-taking features of 54 non-autistic and 26 autistic participants. The extracted speech and turn-taking parameters showed high potential as a diagnostic marker. A linear support vector machine was able to predict the dyad type with 76.2% balanced accuracy (sensitivity: 73.8%, specificity: 78.6%), suggesting that digitally assisted diagnostics could significantly enhance the current clinical diagnostic process due to their objectivity and scalability. In group comparisons on the individual and dyadic level, we found that autistic interaction partners talked slower and in a more monotonous manner than non-autistic interaction partners and that mixed dyads consisting of an autistic and a non-autistic participant had increased periods of silence, and the intensity, i.e. loudness, of their speech was more synchronous.
1. Introduction
Speech as a form of communication is unique to humans. According to Ferdinand de Saussure, it is based on signs combining acoustic forms (the signifier) with meaning (the signified) (). All signifiers can vary in their production to add contextual meaning to their signified. Important speech features like pitch, referring to the tone of speech, intensity, referring to the volume of speech, and articulation rate, referring to the speed of speech, are all influenced by the affective and mental state of the speaker (, ). Therefore, they strongly influence how a certain utterance is perceived: Meaning is not only what we say, but how we say it.
Autism spectrum disorder (ASD) is a neurodevelopmental disorder that entails symptoms regarding communication, social behaviour and behavioural rigidity (). Speech can be completely absent in autistic people. Even in verbal individuals, speech of autistic people differs to that of non-autistic people (, ). One of the most common diagnostic instruments for ASD, the Autism Diagnostic Observation Schedule [ADOS(R); ()], highlights that both changes in prosody and speech rate may indicate ASD, amongst other verbal behaviours.
A recent meta-analysis evaluated the alterations of speech features in ASD (). The authors found that pitch differs between autistic and non-autistic people in terms of increased mean and variance. However, results concerning intensity and speech rate were more equivocal. For both domains, many included studies did not show differences between autistic and non-autistic people, while other studies found effects, though not all of them in the same direction. The meta-analysis did not include studies that investigated variance of intensity over the course of a conversation. In another systematic review (), two studies investigating variance of intensity were mentioned, one of which did not find differences in intensity range (), and the other found decreased standard deviation of intensity (). It is important to note that both Fusaroli et al. () and Asghari et al. () included various modes of speech production, ranging from spontaneous production over narration to social interactions. Additionally, both included all age ranges, so it is possible that not all outcomes apply to adults.
In addition to the importance of speech differences, autistic people report having difficulties with small talk and are perceived as more awkward in conversations (). Since small talk and conversations with strangers are essential for building connections with others, it is important to understand how autistic verbal behaviours differ from non-autistic verbal behaviours in these situations. Reciprocal communication is characterised by a to and fro of speaking and listening. Successful turn-taking not only requires mutual prediction of an upcoming transition point but also a minute concertation of behaviours between interaction partners allowing them to be in sync (). The length of turn-taking gaps can be an estimate of how in sync interaction partners were and is associated with social connection (). If two strangers lose their flow, they tend to feel awkward and try to fill the silence (). A recent study by Ochi et al. () found increased turn-taking gaps and more silence vs. talking as measured by the silence-to-turn ratio (). However, the sample consisted of only male autistic and non-autistic participants, and it is unclear whether the results generalise to people of other genders. Therefore, it is especially important to investigate turn structure in a more general sample to assess the quality of verbal communication.
Finally, the investigation of speech features should be extended to include the temporal fine-tuning within interaction dyads, given the increasing literature showing reduced interactional synchrony in dyads of one autistic and one non-autistic compared to two non-autistic interaction partners [e.g. (); for a review, see ()]. Behavioural synchrony is the product of coordination between interaction partners. This coordination can be achieved by the interaction partners adapting their behaviour to each other. Synchrony of speech features is well documented (); however, research investigating speech synchrony in autistic people is scarce. Ochi et al. () found that non-autistic participants showed more synchrony between the ADOS interviewer’s intensity and their own than autistic participants, but they found no differences regarding synchrony of pitch. Wynn et al. () altered the speed in trial prompts and found that non-autistic adults adapted the speed of their answer in the corresponding trial, while autistic adults and children did not. Both studies show that interpersonal coordination of speech features is a promising avenue to investigate differences in verbal interaction between autistic and non-autistic people.
Additionally, a recent study also used parts of ADOS interviews to investigate classification between autistic and non-autistic children based on synchrony of speech features (). They extracted lexical features and calculated the similarity of the lexical content of the interviews. Machine learning classifiers were able to predict whether a child was diagnosed with ASD with better accuracy when the synchrony measures were added to the model as compared to a model that only included individual speech features. However, in that study, the ADOS was used both for creating the true labels and to extract features for the classification, risking circularity that might artificially inflate accuracies. Therefore, it is vital to assess the performance of classifiers with features extracted from data that is independent from the diagnostic process. In a recent study using automatically extracted interpersonal synchrony of motion quantity and facial expressions, we show that pursuing more naturalistic study designs can yield high classification accuracy of almost 80% (). If these results can be extended to speech and interactional features of verbal communication in adults, this would provide a low-tech and scalable route to assist clinicians with the diagnosis of ASD.
This study design fills the outlined gaps in the literature by extracting speech parameters with an automated pipeline from naturalistic conversations that are independent of the diagnostic assessment to avoid any circularity in the classification procedure. The automated extraction of features increases objectivity, specificity and applicability of the pipeline to a variety of conversational paradigms. The main aim of the current study was (i) to determine the potential of speech coordination as a diagnostic marker for ASD. Additionally, we defined two secondary aims: (ii) to describe individual speech feature differences, and (iii) interactional speech differences that can help explain the classification power. Concerning our main aim (i), we expected that a multivariable prediction model would be able to classify dyad type based on individual speech and dyadic conversational features, thereby offering an exciting possibility for assisting diagnostics of ASD. On the individual level regarding our aim (ii), we expected that autistic and non-autistic individuals would differ in their pitch variance, intensity variance and articulation rate. Additionally, we computed turn-based adaptation of pitch, intensity and articulation rate and expected increased turn-based adaptation in non-autistic compared to autistic individuals. On the dyadic level regarding our aim (iii), we hypothesised that interactional differences would be found in silence-to-turn ratios, turn-taking gaps as well as time-course synchrony of pitch and intensity.
2. Materials and methods
This study is part of a larger project to find diagnostic markers for ASD. The preregistration of the hypotheses regarding aim (ii) and (iii) can be retrieved from OSF.1 Preprocessing was performed using Praat 6.2.09 (), the uhm-o-meter scripts provided by De Jong et al. (, ) and R 4.2.2 () in Rstudio 2022.12.0 (). The Bayesian analysis was performed in R and JASP 0.16.4 (). The machine learning analysis was conducted with the NeuroMiner toolbox 1.1 () implemented in MATLAB R2022b () and Python 3.9.2 All code used to preprocess and analyse the data can be found on GitHub.3 We report our prediction model following the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD) guidelines ().
2.1. Participants
We recruited 35 autistic and 69 non-autistic participants from the general population and the outpatient clinic at the LMU University Hospital Munich by posting flyers at the university and at the hospital as well as distributing them online on social media and mailing lists. Of these participants, 26 autistic (mean age = 34.85 ± 12.01 years, 17 male) and 54 non-autistic (mean age = 30.80 ± 10.42 years, 21 male) participants were analysed (Figure 1). Non-autistic participants were recruited to match the overall gender and age distribution of the autistic sample. This sample is a subset of the sample analysed by Koehler et al. () containing all participants with sufficient audio data quality.
Figure 1
All participants were between 18 and 60 years old, had no current neurological disorder and had an IQ above 70 based on verbal and non-verbal IQ tests (, ). For each autistic participant, an ASD diagnosis (F84.0 or F84.5) according to the ICD-10 () was confirmed by evaluating the diagnostic report. All non-autistic participants had no current or previous psychiatric diagnosis and no intake of psychotropic medication. Autistic and non-autistic participants did not differ credibly in age, verbal [measured with the Mehrfachwahl-Wortschatz-Intelligenztest, MWT-B; ()] or nonverbal IQ [measured with the Culture Fair Intelligence Test, CFT-20-R; ()], but they differed credibly on the Adult Dyspraxia Checklist [ADC; ()], the Autism Quotient [AQ; ()], the Beck’s Depression Inventory [BDI; ()], Self-Monitoring Scale [SMS-short; ()], the Saarbrückener Persönlichkeitsfragebogen [SPF; (), German version of the Interpersonal Reacitivity Index, IRI, ()] and the Toronto Alexithymia Scale [TAS-20; (); see Table 1]. Two autistic participants had a comorbid diagnosis of attention deficit hyperactivity disorder (ADHD), nine of an affective disorder and five of a neurotic stress-related or somatoform disorder. The study was conducted in accordance with the Declaration of Helsinki and approved by the ethics committee of the medical faculty of the LMU. All participants provided written, informed consent and received a monetary compensation for their participation.
Table 1
| Autistic | Non-autistic | Log(BF10) | W | |
|---|---|---|---|---|
| Age | 34.85 ± 12.01 | 30.80 ± 10.42 | −1.028 | 564.50 |
| IQ – nonverbal | 115.35 ± 22.96 | 117.07 ± 15.21 | −1.384 | 716.50 |
| IQ – verbal | 112.12 ± 15.01 | 113.96 ± 16.53 | −1.297 | 743.00 |
| ADC | 50.12 ± 16.06 | 15.57 ± 8.99 | 9.339 | 61.50 |
| AQ | 33.00 ± 8.41 | 14.26 ± 4.55 | 8.863 | 65.50 |
| BDI | 18.35 ± 12.36 | 3.94 ± 3.96 | 8.137 | 139.00 |
| SMS-short | 5.92 ± 2.86 | 9.54 ± 2.98 | 5.080 | 1,136.00 |
| SPF | 36.77 ± 6.62 | 45.43 ± 5.35 | 6.498 | 1,214.00 |
| TAS-20 | 61.27 ± 11.74 | 36.91 ± 7.61 | 8.359 | 77.00 |
Mean and standard deviation of the autistic and non-autistic samples analysed in this study as well as group comparisons performed with Bayesian Mann–Whitney U tests based on 10,000 samples.
Note. ADC, Adult Dyspraxia Scale; AQ, Autism Quotient; BDI, Beck’s Depression Inventory; SMS-short, Self-Monitoring Scale; SPF, Saarbrückener Persönlichkeitsfragebogen; TAS-20, Toronto Alexithymia Scale.
2.2. Experimental procedure
After giving informed consent, blood samples were taken, followed by demographics and the intelligence assessments. Throughout the session, participants completed the above listed questionnaires. They also performed a task assessing emotion recognition [BERT, ()]. In addition, some of the participants took part in a separate study measuring endocrinology and effects of social ostracism.
We paired participants in either mixed dyads consisting of one autistic and one non-autistic participant or non-autistic dyads. Participants were paired based on availability regardless of age and gender. Dyads did not differ in average age or age difference between the interaction partners. However, there was strong evidence in favour of a difference in gender composition (mixed dyads: mean age = 33.15 ± 7.72, mean age difference = 12.69 ± 9.18 [1 to 32 years], 15% female, 35% male and 50% gender-mixed dyads; non-autistic dyads: mean age = 30.18 ± 8.22, mean age difference = 10.64 ± 11.15 [1 to 31 years], 50% female and 50% gender-mixed; for statistical values see Supplementary material S1.1). We did not disclose their interaction partner’s diagnostic status to them. The dyads engaged in two 10-minute long conversations: one about their hobbies and one fun task in which they were asked to plan a menu consisting of food and drinks that they both disliked (). On the one hand, we chose the hobbies task because special interests are a core symptom of ASD (). On the other hand, the meal planning task facilitates a more collaborative interaction and has been shown to promote increased synchrony in non-autistic dyads (). The experimenter left the room during the conversations. After both conversation tasks, participants were asked to rate the quality of their interactions. During the COVID-19 pandemic, testing had to be moved to a different room after nine dyads and a plexiglass was placed between the participants as a health and safety measure. Participants did not wear masks during the conversations and the quality of interactions was rated equal before and after the measures had been put into place ().
We captured participants’ behaviour via multiple channels. The current study focuses on speech coordination captured with one recording device to which two separate microphones were connected (t.Bone earmic 500 with ZoomH4n recorder). The nonverbal communication parameters, body movement captured by a scene camera (Logitech C922), facial expressions captured by two face cameras (Logitech C922), heart rate and electrodermal activity captured by wearables (Empatica E4), as well as the analysis of the blood samples were outside of the scope of the current analysis and published elsewhere (). For more details on the data collection procedure, please consult Supplementary material S1.2.
2.3. Preprocessing
We extracted individual phonetic features for each task and participant using praat () (for more details, see Supplementary material S1.3). We calculated pitch and intensity synchrony with rMEA’s cross-correlation function to calculate windowed cross-lagged correlations (WCLC) using the same window length of 16 s, step size of 8 s and lag of 2 s as Ochi et al. (). We used the uhm-o-meter (, ) to extract turns from conversations, with a turn defined as all speaking instances of one interactant until the end of the speaking instance preceding the next speaking instance of someone else (see Figure 2). For each turn, we calculated turn-taking gap, average pitch, average intensity and number of syllables to calculate articulation rate. Additionally, we used turn-based information to calculate how much each participant adapted their pitch, intensity and articulation rate to the pitch, intensity and articulation rate of the previous turn.
Figure 2
2.4. Comparison of synchrony with pseudosynchrony
We used segment shuffling as described by Moulder et al. (
Table 2
| Mean and SD of values | Mean and SD of pseudo values | Log(BF10) | |
|---|---|---|---|
| Individual adaptation | |||
| Turn-based pitch | 0.121 ± 0.107 | 0.087 ± 0.020 | 4.561 |
| Turn-based intensity | 0.146 ± 0.086 | 0.091 ± 0.017 | 20.977 |
| Turn-based articulation rate | 0.138 ± 0.096 | 0.099 ± 0.019 | 7.017 |
| Dyadic synchrony | |||
| Pitch | 0.197 ± 0.022 | 0.190 ± 0.003 | 1.423 |
| Intensity | 0.368 ± 0.048 | 0.164 ± 0.010 | 104.491 |
Comparison of synchrony and turn-based adaptation values with their corresponding pseudo values.
2.5. Support vector machine for classification
We used a linear L2-regularised L2-loss support vector machine (SVM) as implemented by LIBLINEAR in NeuroMiner to predict each individual’s participation in either a non-autistic or mixed dyad to address our main aim (i). SVMs have not only been applied to classify several psychiatric diagnoses (
Table 3
| Individual | Dyadic |
|---|---|
| Articulation rate | Number of turns |
| Number of pauses | Silence-to-turn ratio |
| Number of syllables | Speech rate |
| Phonation time | Synchrony of intensity |
| Turn-based adaptation of articulation rate | Synchrony of pitch |
| Turn-based adaptation of intensity | Turn-taking gap |
| Turn-based adaptation of pitch | |
| Variance of intensity | |
| Variance of pitch |
List of individual and dyadic features.
Note. All features were entered for the meal planning and the hobbies task separately, resulting in 30 features. Articulation rate refers to the number of syllables per phonation time, while speech rate refers to the number of syllables per total time (phonation time and silence).
2.6. Bayesian analysis
We tested our hypotheses regarding aims (ii) and (iii) using Bayesian repeated-measures ANOVAs as implemented in JASP. Each ANOVA included one within-subjects factor (task: meal planning, hobbies) and one between-subjects factor, either diagnostic status (autistic, non-autistic) or dyad type (mixed, non-autistic). We checked for equality of variance and visually inspected whether the residuals were normally distributed. In the case of violations of these assumptions, we computed a non-parametric alternative and compared the results. We used the Bayes Factor to assess the strength of evidence for or against a model or inclusion of a factor. The Bayes Factor is the ratio of marginal likelihoods, thereby quantifying how much more or less likely one model is than the other. We interpreted the logarithmic Bayes Factor according to Jeffrey’s scheme (
There was a credible difference between the gender composition of the non-autistic and the mixed dyads due to no non-autistic male dyads. Since studies have shown differences between genders with regard to language in ASD (
3. Results
3.1. Performance of support vector machine for classification
Our SVM algorithm was able to distinguish between individuals from a non-autistic and a mixed dyad with 76.2% balanced accuracy on the basis of both individual and dyadic speech and communication features. Specifically, 78.6% of the individuals from a non-autistic dyad were correctly labelled as such (specificity), while 73.8% of the individuals from a mixed dyad were assigned the correct label (sensitivity, see Figure 3). While this model performs significantly above chance levels (p < 0.001; area under the curve: 0.81 [CI 0.72–0.92]; please consult Supplementary material S2 for more details on the SVM classifier), it does not outperform a model trained on synchrony of facial expressions with a balanced accuracy of 79.5% or a stacked model with a balanced accuracy of 77.9% including multiple movement parameters automatically extracted from video recordings of dyadic interactions (
Figure 3

This graph shows for each participant the decision score calculated by the SVM classifier with participants from non-autistic dyads in blue and participants from mixed dyads in green. Filled circles were correctly categorised by the classifier, while empty circles were misclassified.
3.2. Group comparisons on the individual and the dyad level
3.2.1. Individual differences between autistic and non-autistic participants
Autistic participants differed from non-autistic participants in their speech features as evidenced by the results of the Bayesian ANOVAs.
3.2.1.1. Pitch
Pitch variance was best explained by a model including task and diagnostic status but not the interaction of the two [Log(BF10) = 5.612]. The analysis of effects across matched models revealed very strong evidence for the inclusion of task and anecdotal evidence for the inclusion of diagnostic status [task: Log(BFincl) = 4.455; diagnostic status: Log(BFincl) = 1.030]. There was anecdotal evidence against the inclusion of the interaction [task × diagnostic status: Log(BFincl) = −0.510]. However, the Q-Q plot of the residuals revealed deviations from the normal distribution and the variances were not homogeneous. Therefore, we computed a Bayesian Mann–Whitney U test to determine whether the anecdotal evidence in favour of an effect of diagnostic status can be reproduced with a non-parametric test, which was the case [Log(BF10) = 0.888, W = 439.00]. Pitch variance was increased in non-autistic compared to autistic participants.
3.2.1.2. Intensity
The best model describing intensity variance was the full model including the predictors task and diagnostic status as well as their interaction [Log(BF10) = 3.205]. The analysis of effects across matched models revealed that this was mainly driven by the interaction with decisive evidence in favour of the interaction effect and moderate and anecdotal evidence against task and diagnostic status, respectively [task × diagnostic status: Log(BFincl) = 5.163; task: Log(BFincl) = −1.544; diagnostic status: Log(BFincl) = −0.436]. Specifically, while intensity variance of autistic participants was increased in the hobbies condition, the reverse was true for non-autistic participants (see Figure 4).
Figure 4

This graph shows the distribution of individual features in the autistic and non-autistic participants as scatterplots, density plots and box plots. The boxes show the interquartile range and the median, while the whiskers show 1.5 times the interquartile range added to the third and subtracted from the first quartile. In the first row, panel (A) shows pitch variance, panel (B) intensity variance and the panel (C) articulation rate. All three features were increased in non-autistic compared to autistic participants. The second row shows the amount participants adapted their pitch (D), intensity (E) and articulation rate (F) to the previous turn. There were no significant differences in adaption of all three speech factors between autistic and non-autistic participants.
3.2.1.3. Articulation rate
Articulation rate was again best described by the full model including task, diagnostic status and the interaction [Log(BF10) = 6.727], with moderate evidence in favour of including diagnostic status as well as strong evidence in favour of including task and the interaction [task × diagnostic status: Log(BFincl) = 2.517; task: Log(BFincl) = 2.656; diagnostic status: Log(BFincl) = 1.517]. Articulation rate was faster in non-autistic than autistic participants.
3.2.1.4. Turn-based adaptation
Last, the null model outperformed all alternative models with anecdotal evidence in favour of the null model for turn-based adaptation of pitch and intensity (see Supplementary material S4). In the case of adaptation of articulation rate, there was anecdotal evidence in favour of the model including task but no other predictor [Log(BF10) = 1.072] with higher articulation rate in the meal planning condition. Since the residuals were not normally distributed, we performed non-parametric tests which confirmed no effect of diagnostic status on all three adaptation parameters (see Supplementary material S4).
3.2.2. Dyadic differences between non-autistic and mixed dyads
Some interactional features differed between non-autistic and mixed dyads; however, others were comparable in both dyad types (see Figure 5).
Figure 5

This graph shows the distribution of dyadic features for mixed and non-autistic dyads. Panel (A) shows the silence-to-turn ratio which was higher in mixed compared to non-autistic dyads. Panel (B) shows turn-taking gaps which were, on average, longer in mixed dyads. The lower panels show time-course synchrony of pitch (C) and intensity (D) with the latter being higher in mixed dyads.
3.2.2.1. Silence-to-turn ratio
The silence-to-turn ratio was best predicted by the full model including both task and dyad type as well as the interaction [Log(BF10) = 4.141]. A closer look at the analysis of effects across matched models revealed strong evidence in favour of the inclusion of task and moderate evidence in favour of the inclusion of the interaction as well as anecdotal evidence against the inclusion of dyad type as a predictor [task × dyad type: Log(BFincl) = 1.690; task: Log(BFincl) = 2.449; dyad type: Log(BFincl) = 0.049]. This seems to be driven by the increased difference between mixed and non-autistic dyads in the meal planning condition; although, in both conditions the silence-to-turn ratio was smaller in the case of non-autistic dyads.
3.2.2.2. Turn-taking gap
Turn-taking gap was best explained by the model only including the predictor dyad type for which there was anecdotal evidence [Log(BF10) = 0.267]. Similarly, there was anecdotal evidence in favour of including dyad type as well as the interaction of dyad type and task but moderate evidence against including task [task × dyad type: Log(BFincl) = 0.877; task: Log(BFincl) = −1.369; dyad type: Log(BFincl) = 0.279]. Turn-taking gaps tended to be slightly longer in the mixed dyads, especially in the meal planning task.
3.2.2.3. Pitch synchrony
Pitch synchrony, as calculated with WCLC, was best explained by the null model, suggesting that interactants of both dyad types adjusted their pitch to a similar extent to each other (see Supplementary material S5).
3.2.2.4. Intensity synchrony
Nonetheless, non-autistic and mixed dyads differed in their WCLC synchrony of intensity with the best model predicting WLCL synchrony of intensity including both task and dyad type but not the interaction [Log(BF10) = 7.150]. Indeed, there is anecdotal evidence against the inclusion of the interaction, while there is decisive evidence for the inclusion of task and moderate evidence for the inclusion of dyad type [task × dyad type: Log(BFincl) = −0.813; task: Log(BFincl) = 5.567; dyad type: Log(BFincl) = 1.576]. Mixed dyads adjusted their intensity more strongly, with more synchrony in the hobbies condition in both dyad types.
3.2.2.5. Comparison of dyads excluding male dyads
We repeated the analyses of silence-to-turn ratio, turn-taking gap, pitch synchrony and intensity synchrony in a limited sample excluding all male dyads to ensure that the found differences are not driven by differences in gender composition between mixed and non-autistic dyads. For all four parameters, the same model was supported by the evidence as the best model as for the full sample (see Supplementary material S6). Therefore, it is unlikely that the found differences were driven by gender composition.
4. Discussion
Differences in verbal communication are an important symptom of ASD (
Regarding our main research question (i), we are able to present a multivariable prediction model that is able to distinguish between mixed and non-autistic dyads with above 75% of balanced accuracy. Automated extraction of speech and interactional features of verbal conversations offer an exciting new avenue for investigating symptoms as well as assisting the diagnosis of ASD. First, automated extraction increases objectivity and replicability while also providing a more detailed and fine-grained perspective on actual speech differences. This fine-grained perspective could in turn inform intervention by focusing on the specific aspects that differ between autistic and non-autistic conversation partners. Additionally, the current diagnostic procedures are time consuming, and recommendations include a combination of semi-structured interviews and neuropsychological assessments (
The here-presented and other studies (
Concerning research question (ii), our results regarding the speech differences between autistic and non-autistic adults differ from a recent meta-analysis (
Despite several interventions aiming at improving verbal communication skills and turn-taking (71–73), there is little research on differences in interactional features of conversations including autistic people. In this study, we investigated the ratio of silence to turns as well as the duration of the gaps between turns to investigate research question (iii). We found that mixed dyads had a credibly higher ratio of silence to turns, especially when collaboratively planning a meal. This indicates that the amount they were silent was higher, and they were speaking less. This is in line with the findings by Ochi et al. (
It is important to note that studies have shown that some differences in interactions can be reduced or even diminished when autistic individuals are interacting with other autistic people [75–80; for a possible theory explaining this phenomenon see Milton, (81)]. Since this study did not include dyads consisting of two autistic people, it is unclear if the found differences would extend to such a scenario. Future research examining interactional features in verbal communication should investigate possible differences not only between mixed and non-autistic dyads, but should also include comparisons with dyads consisting of two autistic interaction partners.
In addition to interactional features of verbal conversations, we also assessed synchrony and turn-based adaptation of speech features between two interaction partners in a dyad. We found no difference in turn-based adaptation between autistic and non-autistic adults, meaning that the extent to which they adapted their pitch, intensity and articulation rate to the previous turn was comparable in both groups. Similarly, we also did not find any differences between time-course synchrony of pitch between mixed and non-autistic dyads. However, we found that time-course synchrony of intensity was higher in mixed dyads than in non-autistic dyads. This is in contrast to Ochi et al. (
Despite the insights this study offers, it is still unclear how context influences speech production with respect to ASD. We aimed for a naturalistic conversation setting with one common (hobbies) and one uncommon (meal planning) conversation topic. Other studies have opted to focus on a more controlled speech production by pairing participants with a trained diagnostician (
In this study, we investigated the potential of speech and interactional features of verbal communication for digitally assisted diagnostics. We used automatic feature extraction on two naturalistic 10-minute conversations between either two non-autistic strangers (non-autistic dyad) or one autistic and one non-autistic stranger (mixed dyad). We were able to classify between individuals from a non-autistic vs. from a mixed dyad based on these features with high accuracy which offers a low-tech, economic and scalable option for diagnostic classification. Additionally, we have shown differences in pitch and intensity variation as well as articulation rate between autistic and non-autistic adults and differences in silence-to-turn ratio, turn-taking gaps and time-course synchrony of intensity between non-autistic and mixed dyads. This study shows the potential of verbal markers for diagnostic classification of ASD and suggests multiple relevant features showing differences between autistic and non-autistic adults.
Statements
Data availability statement
The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.
Ethics statement
The studies involving humans were approved by Ethikkommission der Medizinischen Fakultät der Ludwig-Maximilians-Universität München. The studies were conducted in accordance with the local legislation and institutional requirements. The participants provided their written informed consent to participate in this study.
Author contributions
IP: Conceptualization, Formal analysis, Investigation, Methodology, Software, Supervision, Visualization, Writing – original draft, Writing – review & editing. JK: Conceptualization, Data curation, Funding acquisition, Investigation, Methodology, Project–administration, Writing – review & editing. AN: Conceptualization, Data curation, Investigation, Methodology, Project–administration, Writing – review & editing. NK: Conceptualization, Methodology, Resources, Supervision, Writing – review & editing. CF-W: Conceptualization, Funding acquisition, Investigation, Methodology, Resources, Supervision, Writing – review & editing.
Funding
The author(s) declare financial support was received for the research, authorship, and/or publication of this article. This work was supported by Stiftung Irene (PhD scholarship awarded to JK) and the German Research Council (grant numbers 876/3–1 and FA 876/5–1 awarded to CF-W).
Acknowledgments
We want to thank Elsa Sangaran for their part in writing the preregistration and inspecting the data. We also want to show appreciations to the interns in our group who aurally and visually inspected the raw and the preprocessed data (in alphabetical order): Alena Holy, Christian Aldenhoff and Sophie Herke. Additionally, we want to thank Stephanie Fischer and Johanna Späth for their contribution to the data collection and Mark Sen Dong for providing scripts to create cross validation structures. Last, we want to express our thanks to Francesco Cangemi for their input on the microphone setup and the preprocessing parameters.
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. The author(s) declared that they were an editorial board member of Frontiers, at the time of submission. This had no impact on the peer review process and the final decision.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Supplementary material
The Supplementary material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fpsyt.2023.1257569/full#supplementary-material
References
1.
LepschyG. F. de Saussure, Course in general linguistics, translated and annotated by Roy Harris. London: Duckworth, 1983. J Linguist. (1985) 21:250–4. doi: 10.1017/S0022226700010185
2.
PereiraCWatsonC. Some acoustic characteristics of emotion. 5th Int Conf Spok Lang Process ICSLP 1998. (1998). doi: 10.21437/icslp.1998-148
3.
WhitesideSP. Acoustic characteristics of vocal emotions simulated by actors. Percept Mot Skills. (1999) 89:1195–208. doi: 10.2466/pms.1999.89.3f.1195
4.
World Health Organization. (2019). International Classification of Diseases, Eleventh Revision (ICD-11). Available at: https://icd.who.int/browse11.
5.
AsghariSZFarashiSBashirianSJenabiE. Distinctive prosodic features of people with autism spectrum disorder: a systematic review and meta-analysis study. Sci Rep. (2021) 11:23093–17. doi: 10.1038/s41598-021-02487-6
6.
PickettEPullaraOO’GradyJGordonB. Speech acquisition in older nonverbal individuals with autism: a review of features, methods, and prognosis. Cogn Behav Neurol. (2009) 22:1–21. doi: 10.1097/WNN.0b013e318190d185
7.
LordCRutterMGoodeSHeemsbergenJJordanHMawhoodLet al. Austism diagnostic observation schedule: a standardized observation of communicative and social behavior. J Autism Dev Disord. (1989) 19:185–212. doi: 10.1007/BF02211841
8.
FusaroliRLambrechtsABangDBowlerDMGaiggSB. Is voice a marker for autism spectrum disorder? A systematic review and meta-analysis. Autism Res. (2017) 10:384–407. doi: 10.1002/aur.1678
9.
GrossmanRBBemisRHPlesa SkwererDTager-FlusbergH. Lexical and affective prosody in children with high functioning autism. J Speech Lang Hear Res. (2010) 53:778–93. doi: 10.1044/1092-4388(2009/08-0127)
10.
ScharfsteinLABeidelDCSimsVKRendon FinnellL. Social skills deficits and vocal characteristics of children with social phobia or asperger’s disorder: a comparative study. J Abnorm Child Psychol. (2011) 39:865–75. doi: 10.1007/s10802-011-9498-2
11.
BoneDBlackMPRamakrishnaAGrossmanRNarayananS. Acoustic-prosodic correlates of ‘awkward’ prosody in story retellings from adolescents with autism. Proc Annu Conf Int Speech Commun Assoc INTERSPEECH. (2015):1616–20. doi: 10.21437/interspeech.2015-374
12.
EliasRMuskettAEWhiteSW. Educator perspectives on the postsecondary transition difficulties of students with autism. Autism. (2019) 23:260–4. doi: 10.1177/1362361317726246
13.
GarrelsV. Getting good at small talk: student-directed learning of social conversation skills. Eur J Spec Needs Educ. (2019) 34:393–402. doi: 10.1080/08856257.2018.1458472
14.
GrossmanRB. Judgments of social awkwardness from brief exposure to children with and without high-functioning autism. Autism. (2015) 19:580–7. doi: 10.1177/1362361314536937
15.
TomprouMKimYJChikersalPWoolleyAWDabbishLA. Speaking out of turn: how video conferencing reduces vocal synchrony and collective intelligence. PloS One. (2021) 16:e0247655–14. doi: 10.1371/journal.pone.0247655
16.
TempletonEMChangLJReynoldsEALeBeaumontMDCWheatleyTPCone LebeaumontMDet al. Fast response times signal social connection in conversation. PNAS. (2022) 119:1–8. doi: 10.1073/pnas.2116915119
17.
McLaughlinMLCodyMJ. Awkward silences: behavioral antecedents and consequences of the conversational lapse. Hum Commun Res. (1982) 8:299–316. doi: 10.1111/j.1468-2958.1982.tb00669.x
18.
OchiKOnoNOwadaKKojimaMKurodaMSagayamaSet al. Quantification of speech and synchrony in the conversation of adults with autism spectrum disorder. PloS One. (2019) 14:e0225377–22. doi: 10.1371/journal.pone.0225377
19.
NelsonA. M. (2020). Investigating the acoustic, prosodic, and interactional speech characteristics of autistic traits: An automated speech analysis in typically-developed adults.
20.
GeorgescuALKoerogluSHamiltonAFCVogeleyKFalter-WagnerCMTschacherW. Reduced nonverbal interpersonal synchrony in autism spectrum disorder independent of partner diagnosis: a motion energy study. Mol Autism. (2020) 11:11–4. doi: 10.1186/s13229-019-0305-1
21.
McNaughtonKARedcayE. Interpersonal synchrony in autism. Curr Psychiatry Rep. (2020) 22:1–11. doi: 10.1007/s11920-020-1135-8
22.
GregorySWebsterSHuangG. Voice pitch and amplitude convergence as a metric of quality in dyadic interviews. Lang Commun. (1993) 13:195–217. doi: 10.1016/0271-5309(93)90026-J
23.
NataleM. Convergence of mean vocal intensity in dyadic communication as a function of social desirability. J Pers Soc Psychol. (1975) 32:790–804. doi: 10.1037/0022-3514.32.5.790
24.
StreetRL. Speech convergence and speech evaluation in fact-finding interviews. Hum Commun Res. (1984) 11:139–69. doi: 10.1111/j.1468-2958.1984.tb00043.x
25.
WardA.LitmanD. (2007). Measuring convergence and priming in tutorial dialog. Univ. Pittsburgh. Available at: http://scholar.google.com/scholar?hl=en&btnG=Search&q=intitle:Measuring+Convergence+and+Priming+in+Tutorial+Dialog#0.
26.
WynnCJBorrieSASellersTP. Speech rate entrainment in children and adults with and without autism spectrum disorder. Am J Speech-Language Pathol. (2018) 27:965–74. doi: 10.1044/2018_AJSLP-17-0134
27.
LahiriRNasirMKumarMKimSHBishopSLordCet al. Interpersonal synchrony across vocal and lexical modalities in interactions involving children with autism spectrum disorder. JASA Express Lett. (2022) 2:095202. doi: 10.1121/10.0013421
28.
KoehlerJCDongMSNelsonAMFischerSSpäthJPlankISet al. Machine learning classification of autism Spectrum disorder based on reciprocity in naturalistic social interactions. medRxiv. (2022):22283571. doi: 10.1101/2022.12.20.22283571
29.
BoersmaP.WeeninkD. (2022). Praat: doing phonetics by computer. Available at: http://www.praat.org/.
30.
De JongNHPacillyJHeerenW. PRAAT scripts to measure speed fluency and breakdown fluency in speech automatically. Assess Educ Princ Policy Pract. (2021) 28:456–76. doi: 10.1080/0969594X.2021.1951162
31.
De JongN. H.PacillyJ.HeerenW. (2021). Uhm-o-meter [computer software]. Available at: https://sites.google.com/view/uhm-o-meter/home.
32.
R Core Team. (2021). R: A Language and Environment for Statistical Computing. Available at: https://www.r-project.org.
33.
RStudio Team. (2020). RStudio: Integrated Development Environment for R. Available at: http://www.rstudio.com/.
34.
JASP Team. (2022). JASP (Version 0.16.4)[Computer software]. Available at: https://jasp-stats.org/.
35.
KoutsoulerisN.VetterC.WiegandA. (2022). Neurominer [computer software]. Available at: https://github.com/neurominer-git/NeuroMiner_1.1.
36.
The Mathworks Inc. (2022). MATLAB version: 9.13.0 (R2022b), Natick, Massachusetts: The MathWorks Inc. https://www.mathworks.com
37.
CollinsGSReitsmaJBAltmanDGMoonsKGM. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): the TRIPOD statement. Ann Intern Med. (2015) 162:55–63. doi: 10.7326/M14-0697
38.
LehrlS. (1999). Mehrfachwahl-Wortschatz-Intelligenztest: MWT-B. Spitta.
39.
WeißR. (2006). CFT 20-R: grundintelligenztest skala 2-revision. Hogrefe.
40.
World Health Organization. The ICD-10 classification of mental and behavioural disorders: Diagnostic criteria for research. Geneva: World Health Organization, (1993).
41.
KirbyAEdwardsLSugdenDRosenblumS. The development and standardization of the adult developmental co-ordination disorders/dyspraxia checklist (ADC). Res Dev Disabil. (2010) 31:131–9. doi: 10.1016/j.ridd.2009.08.010
42.
Baron-CohenSWheelwrightSSkinnerRMartinJClubleyE. The autism-Spectrum quotient (AQ): evidence from Asperger syndrome/high-functioning autism, males and females, scientists and mathematicians. J Autism Dev Disord. (2001) 31:5–17. doi: 10.1023/A:1005653411471
43.
HautzingerM.BailerM.WorallH.KellerF. (1994). Beck-depressions-inventar (BDI). Huber.
44.
GrafA. A German version of the self-monitoring scale. Zeitschrift fur Arbeits- und Organ. (2004) 48:109–21. doi: 10.1026/0932-4089.48.3.109
45.
PaulusC. (2009). Der Saarbrücker Persönlichkeitsfragebogen SPF (IRI) zur Messung von Empathie: Psychometrische Evaluation der deutschen Version des Interpersonal Reactivity Index. Available at: http://hdl.handle.net/20.500.11780/3343.
46.
DavisM. H. (1980). Interpersonal reactivity index. Available at: 10.1037/t01093-000
47.
PoppKSchäferRSchneiderCBrählerEDeckerOHardtJet al. Faktorstruktur und Reliabilität der Toronto-Alexithymie-Skala (TAS-20) in der deutschen Bevölkerung. Psychother Med Psychol. (2008) 58:208–14. doi: 10.1055/s-2007-986196
48.
DrimallaH.BaskowI.BehniaBet al (2019). Imitation and recognition of facial emotions in autism: a computer vision approach. Molecular Autism. (2021) 12:27. doi: 10.1186/s13229-021-00430-0
49.
TschacherWReesGMRamseyerF. Nonverbal synchrony and affect in dyadic interactions. Front Psychol. (2014) 5:1–13. doi: 10.3389/fpsyg.2014.01323
50.
MoulderRGBokerSMRamseyerFTschacherW. Determining synchrony between behavioral time series: an application of surrogate data generation for establishing falsifiable null-hypotheses. Psychol Methods. (2018) 23:757–73. doi: 10.1037/met0000172
51.
DwyerDBFalkaiPKoutsoulerisN. Machine learning approaches for clinical psychology and psychiatry. Annu Rev Clin Psychol. (2018) 14:91–118. doi: 10.1146/annurev-clinpsy-032816-045037
52.
OrrùGPettersson-YeoWMarquandAFSartoriGMechelliA. Using support vector machine to identify imaging biomarkers of neurological and psychiatric disease: a critical review. Neurosci Biobehav Rev. (2012) 36:1140–52. doi: 10.1016/j.neubiorev.2012.01.004
53.
GeorgescuALKoehlerJCWeiskeJVogeleyKKoutsoulerisNFalter-WagnerCM. Machine learning to study social interaction difficulties in ASD. Front Robot AI. (2019) 6:1–7. doi: 10.3389/frobt.2019.00132
54.
KoehlerJCGeorgescuALWeiskeJSpangemacherMBurghofLFalkaiPet al. Brief report: specificity of interpersonal synchrony deficits to autism Spectrum disorder and its potential for digitally assisted diagnostics. J Autism Dev Disord. (2022) 52:3718–26. doi: 10.1007/s10803-021-05194-3
55.
FanR.-E.ChangK.-W.HsiehC.-J.WangX.-R.LinC.-J. (2022). LIBLINEAR: a library for large linear classification. Available at: https://www.csie.ntu.edu.tw/~cjlin/papers/liblinear.pdf.
56.
PavlouMAmblerGSeamanSDe IorioMOmarRZ. Review and evaluation of penalised regression methods for risk prediction in low-dimensional data with few events. Stat Med. (2016) 35:1159–77. doi: 10.1002/sim.6782
57.
CortesCVapnikVN. Support-vector networks. Mach Learn. (1995) 20:273–97. doi: 10.1007/BF00994018
58.
VapnikVN. An overview of statistical learning theory. IEEE Trans Neural Netw. (1999) 10:988–99. doi: 10.1109/72.788640
59.
VapnikVN. The nature of statistical learning theory. 2nd ed. New York: Springer (2000).
60.
Goss-SampsonM. Bayesian inference in JASP (2020):1–120. doi: 10.17605/OSF.IO/CKNXM,
61.
BoorseJColaMPlateSYankowitzLPandeyJSchultzRTet al. Linguistic markers of autism in girls: evidence of a ‘blended phenotype’ during storytelling. Mol Autism. (2019) 10:1–12. doi: 10.1186/s13229-019-0268-2
62.
SturrockAAdamsCFreedJ. A subtle profile with a significant impact: language and communication difficulties for autistic females without intellectual disability. Front Psychol. (2021) 12:1–9. doi: 10.3389/fpsyg.2021.621742
63.
SturrockAChiltonHFoyKFreedJAdamsC. In their own words: the impact of subtle language and communication difficulties as described by autistic girls and boys without intellectual disability. Autism. (2022) 26:332–45. doi: 10.1177/13623613211002047
64.
ZwaigenbaumLPennerM. Autism spectrum disorder: advances in diagnosis and evaluation. BMJ. (2018) 361:k1674–16. doi: 10.1136/bmj.k1674
65.
MatsonJLKozlowskiAM. The increasing prevalence of autism spectrum disorders. Res Autism Spectr Disord. (2011) 5:418–25. doi: 10.1016/j.rasd.2010.06.004
66.
RuzichEAllisonCSmithPWatsonPAuyeungBRingHet al. Measuring autistic traits in the general population: a systematic review of the autism-Spectrum quotient (AQ) in a nonclinical population sample of 6,900 typical adult males and females. Mol Autism. (2015) 6:1–12. doi: 10.1186/2040-2392-6-2
67.
KalandCSwertsMKrahmerE. Accounting for the listener: comparing the production of contrastive intonation in typically-developing speakers and speakers with autism. J Acoust Soc Am. (2013) 134:2182–96. doi: 10.1121/1.4816544
68.
ChanKKLToCKS. Do individuals with high-functioning autism who speak a tone language show intonation deficits?J Autism Dev Disord. (2016) 46:1784–92. doi: 10.1007/s10803-016-2709-5
69.
DePapeAMRChenAHallGBCTrainorLJ. Use of prosody and information structure in high functioning adults with autism in relation to language ability. Front Psychol. (2012) 3:1–13. doi: 10.3389/fpsyg.2012.00072
70.
HubbardDJFasoDJAssmannPFSassonNJ. Production and perception of emotional prosody by adults with autism spectrum disorder. Autism Res. (2017) 10:1991–2001. doi: 10.1002/aur.1847
71.
BambaraLMThomasAChovanesJColeCL. Peer-mediated intervention: enhancing the social conversational skills of adolescents with autism Spectrum disorder. Teach Except Child. (2018) 51:7–17. doi: 10.1177/0040059918775057
72.
RiethSRStahmerACSuhrheinrichJSchreibmanLKennedyJRossB. Identifying critical elements of treatment: examining the use of turn taking in autism intervention. Focus Autism Other Dev Disabl. (2014) 29:168–79. doi: 10.1177/1088357613513792
73.
ThirumanickamARaghavendraPMcMillanJMvan SteenbruggeW. Effectiveness of video-based modelling to facilitate conversational turn taking of adolescents with autism spectrum disorder who use AAC. AAC Augment Altern Commun. (2018) 34:311–22. doi: 10.1080/07434618.2018.1523948
74.
BoneDBishopSGuptaRLeeSNarayananSS. Acoustic-prosodic and turn-taking features in interactions with children with neurodevelopmental disorders. Proc Annu Conf Int Speech Commun Assoc INTERSPEECH. (2016):1185–9. doi: 10.21437/Interspeech.2016-1073
75.
HeasmanBGillespieA. Neurodivergent intersubjectivity: distinctive features of how autistic people create shared understanding. Autism. (2019) 23:910–21.
76.
CromptonCJRoparDEvans-WilliamsCVFlynnEGFletcher-WatsonS. Autistic peer-to-peer information transfer is highly effective. Autism. (2020) 24:1704–12.
77.
CromptonCJHallettSRoparDFlynnEFletcher-WatsonS. ‘I never realised everybody felt as happy as I do when I am around autistic people’: A thematic analysis of autistic adults’ relationships with autistic and neurotypical friends and family. Autism. (2020) 24:1438–48.
78.
CromptonCJSharpMAxbeyHFletcher-WatsonSFlynnEGRoparD. Neurotype-matching, but not being autistic, influences self and observer ratings of interpersonal rapport. Front Psychol. (2020) 11:2961.
79.
MorrisonKEDeBrabanderKMJonesDRFasoDJAckermanRASassonNJ. Outcomes of real-world social interaction for autistic adults paired with autistic compared to typically developing partners. Autism. (2020) 24:1067–80.
80.
AlkireDMcNaughtonKAYargerHAShariqDRedcayE. Theory of mind in naturalistic conversations between autistic and typically developing children and adolescents. Autism. (2023) 27:472–88.
81.
MiltonDE. On the ontological status of autism: The ‘double empathy problem’. Disability and society. (2012) 27:883–7.
Summary
Keywords
speech, turn-taking, diagnostic classification, autism, prediction model, conversation
Citation
Plank IS, Koehler JC, Nelson AM, Koutsouleris N and Falter-Wagner CM (2023) Automated extraction of speech and turn-taking parameters in autism allows for diagnostic classification using a multivariable prediction model. Front. Psychiatry 14:1257569. doi: 10.3389/fpsyt.2023.1257569
Received
21 July 2023
Accepted
20 October 2023
Published
06 November 2023
Volume
14 - 2023
Edited by
Lawrence Fung, Stanford University, United States
Reviewed by
Philippine Geelhand, Université libre de Bruxelles, Belgium; Kenji J. Tsuchiya, Hamamatsu University School of Medicine, Japan
Updates

Check for updates
Copyright
© 2023 Plank, Koehler, Nelson, Koutsouleris and Falter-Wagner.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: I. S. Plank, irene.plank@med.uni-muenchen.de
†ORCID: I. S. Plank, https://orcid.org/0000-0002-9395-0894
J. C. Koehler, https://orcid.org/0000-0001-7216-6075
A. M. Nelson, https://orcid.org/0000-0002-7453-2187
N. Koutsouleris, https://orcid.org/0000-0001-6825-6262
C. M. Falter-Wagner https://orcid.org/0000-0002-5574-8919
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.