Abstract
A foundational aspect of social interaction is the phenomenon of joint attention, typically defined as a shared focus between two or more interlocutors on a particular subject, object or event. While joint attention research in the fields of Cognitive Linguistics and psycholinguistics has predominantly relied on task-based interactions in static environments, studies in the tradition of Interactional Linguistics and Conversation Analysis have focused on the sequential-interactional organization of joint attention in unrestrained interaction but have often been limited to smaller datasets. Moreover, existing work has mainly adopted either quantitative methods (in cognitively oriented research) or qualitative approaches (in interactionally oriented studies), whereas mixed-methods analyses of joint attention remain scarce. To address these methodological research gaps, this study draws on a substantial eye-tracking dataset of naturally occurring interactions in a mobile setting (i.c. joint museum visits) and adopts a mixed-methods approach. On the one hand, we apply an established quantitative method for examining gaze behavior in controlled environments, namely cross-recurrence quantification analysis, to the novel context of mobile interactions. On the other hand, we integrate this quantitative perspective with a multimodal conversation analysis to explore the advantages and limitations of combining methods in joint attention research. Specifically, we also propose the idea of a close-recurrence analysis as a visualization technique that repurposes the central logic of cross-recurrence quantification for qualitative analytic work.
1 Introduction
Ever since seminal work on eye gaze as a semiotic resource, gaze has been known to play a crucial role in social interaction (Zima and Stukenbrock, 2025). To systematically analyze eye gaze in interaction, research in the past two decades has typically employed eye-tracking technologies, which function as a valuable tool to gather detailed information on interactants' gaze behavior (see , for overviews). In particular, eye-tracking techniques can be divided into remote and mobile eye-tracking. On the one hand, remote eye-tracking refers to screen-based eye-tracking systems, used to analyze lab-based interactions where the participants usually perform interactive tasks on a computer, which records the participants' gaze behaviors (cf. the dual eye-tracking paradigm, involving the specific set-up of two remote eye-tracking systems to analyze collaborative tasks, see for an overview). On the other hand, mobile eye-tracking involves wearable eye-tracking systems, which typically take the form of a pair of glasses. This allows for the study of gaze behavior in maximally unobtrusive conditions, with participants engaging in spontaneous interactions while sitting, standing, walking, driving, etc. (see Zima and Stukenbrock, 2025 for an overview). Importantly for the study of joint attention, all interlocutors can be tracked simultaneously (as a form of multifocal eye-tracking, ), providing fine-grained information on the gaze behavior of all participants at each point in time. Previous eye-tracking research on joint attention is mainly situated in two fields of research: Cognitive Linguistics and psycholinguistics on the one hand and Conversation Analysis and Interactional Linguistics on the other hand (see , for overviews). In this introduction, we focus on the datasets and methods that these research strands have typically employed to analyze the phenomenon of joint attention, drawing on either remote or mobile eye-tracking techniques.
Previous studies adopting a cognitive perspective typically define joint attention as a shared focus between two or more individuals on a particular subject, object or event (Tomasello and Carpenter, 2007). In terms of analytical focus, cognitive research mostly examines the role of eye gaze in relation to joint attention, for example exploring mutual gaze (direct eye contact between two interactants) and shared gaze (two or more interactants looking at the same object of attention) as important resources for grounding, reference tracking, etc. (e.g., Neider et al., 2010). While the interactants' gaze behavior can thus indicate a moment of joint attention, eye gaze may also play an integral part in getting to that point. A well-known phenomenon in this regard is the gaze cueing effect, referring to speakers guiding the gaze direction of their addressees by looking at, and thus cueing, a particular target (McKay et al., 2021). Considering the temporal dimension of eye gaze in interaction, multiple studies have also focused on gaze synchronization, examining if interlocutors (start to) look at each other or a given object simultaneously or with a given time lag (e.g., ; Oben, 2015; Richardson et al., 2007). As such, previous research (e.g., ; Richardson et al., 2009; Vrzakova et al., 2019) has demonstrated that there is a fine-grained systematic coupling between interlocutors’ gaze fixations as interactants tend to look at the same thing at the same time, or at least with a systematic time lag between the areas of interest involved.
To date, cognitively oriented research has mostly been conducted on highly controlled, task-based interactions within lab settings, which typically involve remote eye-trackers and for example puzzle or matching games (; ; ; Richardson et al., 2007). More recent studies have also analyzed less mediated face-to-face interactions with mobile eye-trackers, but these conversations typically still take place within a static lab environment, involving some form of instructional task (e.g., brainstorming in or tutoring in Shi and Stickler, 2021). In the analysis of these datasets, this line of research typically uses quantitative research methods, including various types of distribution and correlation measures. Studies specifically interested in gaze synchronization often rely on a technique called cross-recurrence quantification analysis (CRQA) (see e.g., ; Richardson and Dale, 2005; Richardson et al., 2009). In particular, a cross-recurrence analysis is “a type of correlation analysis that looks for a time lag at which the overlap between two time-series is maximal” (Oben, 2015, p. 137). This allows the analyst to check whether events typically occur simultaneously, with a given time lag, or completely unrelated from one another (see ; Xu et al., 2020 for overviews on research using cross-recurrence techniques to analyze different kinds of behavior matching).
In addition to this cognitively oriented line of research on joint attention, a relatively recent research strand has started to approach the phenomenon from an interactional perspective, defining joint attention as “an interactional accomplishment that involves two (or more) participants who mutually coordinate to establish a triadic relationship with an object or event” (Stukenbrock and Balantani, 2025, p. 248). As such, this line of research focuses on the sequential-interactional organization of joint attention, highlighting the role of various semiotic resources. Previous research characterizes joint attention as an embodied choreography (Tulbert and Goodwin, 2011) and emphasizes that joint attentional processes in interaction draw on both verbal and bodily-visual resources, including discourse particles, (pointing) gestures, eye gaze and the handling of objects (e.g., ; Stukenbrock, 2020; Stukenbrock and Balantani, 2025). Specifically regarding eye gaze, Stukenbrock (2020) has also shown the importance of speaker gaze following (i.e., the addressee following the speaker's gaze direction) and addressee gaze monitoring (i.e., the speaker monitoring the gaze direction of the addressee) in the organization of joint attention. As such, both concepts can be considered the interactionally-embedded counterpart of gaze cueing in cognitive fields of research.
In terms of datasets, the use of eye-tracking technology in interactionally oriented studies is more recent in comparison to cognitively oriented joint attention research. To date, most studies in Interactional Linguistics and Conversation Analysis still rely exclusively on video recordings, which do not allow fine-grained analyses of gaze behavior (Zima et al., 2025). However, there is a growing body of research on gaze, reference and joint attention that uses mobile eye-tracking techniques to capture gaze behavior in various kinds of real life, walking-and-talking interactions, including nature walks (; ; ; ), museum visits (; ; ; Stukenbrock, 2020, 2023; Stukenbrock and Balantani, 2025), furniture building (), visits to a farmer's market (Stukenbrock, 2018a, b, 2020; Stukenbrock and Dao, 2019) and searching for a book at the library (Stukenbrock, 2018a, b). Notably, most of these datasets are limited in terms of size and number of participants (except for the corpus in ). The analysis of these datasets is thus typically situated in the paradigm of Conversation Analysis, involving a qualitative approach from a multimodal perspective.
A key insight from these interactional studies is that joint attention is often initiated with a multimodal summons-answer sequence (; Stukenbrock, 2018b, 2020). Concretely, Stukenbrock and Balantani (2025) describe the organization of joint attention as unfolding over three sequential positions. First, a speaker produces a summons to share attention, which functions as a request for the addressee's gaze. In general, this summons typically includes a demonstrative and an embodied pointing device (; Stukenbrock, 2020). In particular, the summons may also include perceptual directives such as ‘look’ (; ) or addressee gaze monitoring (Stukenbrock, 2020) and take shape in the form of a ‘noticing’ drawing the addressee's attention to the referential situation (Stukenbrock, 2023; Stukenbrock and Dao, 2019). After the summons, the addressee typically produces an answer that involves an embodied re-orientation to the speaker and a gaze shift towards the relevant target or referent (Stukenbrock and Balantani, 2025). Finally, the interlocutors' joint focus of attention is then acknowledged through a display of shared perception or a documentation of understanding (Stukenbrock and Balantani, 2025). In this paper, we conceptualize such summons-answer sequences as a form of leader-followership in the organization of joint attention in interaction, with the summoner sequentially initiating the joint attentional process and the addressee multimodally following the summoner's gaze direction, pointing gestures and/or verbal directives.
While cognitively oriented research on joint attention has predominantly relied on quantitative methods and interactional studies have primarily drawn on qualitative approaches, a few eye-tracking studies on gaze behavior in interaction also adopt a mixed-methods approach (see for a methodological overview). In particular, these studies have considered the role of eye gaze in, for example, turn-taking (), word searches (), question-response sequences (), irony () and overlap resolution (Zima et al., 2019) and have thereby demonstrated the benefits of letting quantitative and qualitative methods feed into each other in eye gaze research. As eye-tracking tools provide detailed and extensive data that can be examined from several analytical perspectives, Oben et al. (2025) also argue that a combined qualitative-quantitative approach may nuance existing results or lead to new insights. However, mixed-methods analyses in the study of eye gaze in interaction remain scarce, particularly for the phenomenon of joint attention (see, however, Oben et al., 2025 for a mixed-methods study on gaze synchronization).
In conclusion, there are a number of research gaps in the current literature on eye-tracking and joint attention. First, quantitative research on joint attention is mostly limited to controlled, lab-based interactions (which reduces the ecological validity of the results). Second, while there is a growing body of qualitative studies on the organization of joint attention in naturally occurring interactions, these studies remain limited to relatively small datasets. Third, mixed-methods research on joint attention remains scarce, although recent studies have emphasized the benefits of mixed-methods approaches for the study of gaze in interaction. To address these research gaps, this study employs a large dataset of real-life, walking-and-talking interactions, enabling both quantitative and qualitative research methods to study the phenomenon of joint attention. Concretely, this study has two methodological research objectives: (i) exploring the use of cross-recurrence quantification analysis beyond the confinements of the lab, in the novel context of mobile interactions and (ii) examining possible avenues for cross-pollination between CRQA and qualitative methods of interaction analysis.
2 Materials and methods
2.1 Dataset
To pursue our research objectives, we employ a dataset of joint museum visits in this study, as the context of a museum exhibition is at the same time naturalistic (i.e., participants are free to navigate their routes and decide which paintings or installations to attend to) and partly controlled (i.e., most visitors encounter the same artworks). As such, this allows for a quantitative approach, comparing joint attentional processes across participants and for longer stretches of time, as well as a qualitative approach to perform fine-grained analyses of specific fragments. Moreover, previous research has explicitly presented museum visits as an activity where joint attention is strongly invited, claiming that visiting a museum together calls for sharing attention on the objects that are displayed as a way of enhancing togetherness (Stukenbrock and Balantani, 2025; Vom Lehn, 2013). In line with this, the majority of existing qualitative studies on joint attention also analyze interactions during museum exhibition visits (see references to Anja Stukenbrock's work above).
Concretely, the dataset in this study consists of a large audio-visual corpus of dyadic interactions (25 dyads, 50 participants, 1,000 min) during joint museum visits. The participants (all native speakers of Dutch aged 18–75) visited the Turning Heads exhibition at the Royal Museum of Fine Arts in Antwerp, after completing an informed consent procedure approved by the KU Leuven Social and Societal Ethics Committee (file nr. G-2023-7074). They were asked to visit the exhibition as they normally would and did not receive any specific instructions, tasks or timing restrictions. Regarding the recording set-up, all participants wore mobile eye-tracking glasses (Tobii Glasses 3) and, in addition, we used an external GoPro camera to record the participants' embodied behavior from a follower's perspective. The recordings were synchronized into one split-screen video using Adobe Premiere Pro (see Figure 1 below) and were transcribed and annotated in ELAN (Wittenburg et al., 2006), according to the conventions for verbal transcription and, for selected fragments, the Mondada (2018) conventions for multimodal annotation. For this study, we selected the data from 15 visitor pairs (i.e., 30 participants, with various interpersonal relations1) in two exhibition rooms (including 29 artworks in total), which amounts to 4.5 h of data. In terms of duration, this does not significantly exceed the datasets used in previous qualitative studies on joint attention (see above), but, importantly, this dataset includes a larger number of participants, which allows for a quantitative as well as a qualitative analysis.
Figure 1
2.2 Quantitative method of analysis
To prepare the dataset for a cross-recurrence quantification analysis (CRQA), we annotated the participants' gaze fixations on the different artworks, labels (typically including the artwork's title, the artist's name and a date), information texts (longer texts that provide background information on the artwork and/or artist) and their co-participant, giving each area of interest (AOI) a distinct annotation value. In particular, we adopted a minimal fixation duration of 120 ms (in line with previous studies such as Oben, 2015; Vertegaal et al., 2001) and we included gaze shifts towards a certain AOI as a part of the fixation on that AOI. In a next step, the gaze annotations were filtered to only include fixations overlapping with speech, given our interest in interactional effects or behavioral matching rather than contextual effects due to the spatial constraints of a museum exhibition. In doing so, we also included gaze fixations during the 500 ms surrounding each utterance to account for minimal pauses in the interaction. In addition, we included the whole annotation for gaze fixations at the start or end of a sequence, to capture interactions in response to a certain gaze focus or vice versa. Following these additional selection steps, 1.3 h of data remain for the ensuing CRQA. To check for the quality and consistency of the gaze annotations, we also conducted an inter-coder agreement (ICA) analysis based on the selected gaze annotations of two visitor pairs (i.c. dyads 39_40 and 58_59, infra). This ICA analysis rendered a Cohen's Kappa score of 0.96, which is considered an almost perfect agreement according to .
As explained above, CRQA is a type of correlation analysis that identifies the time lag at which two time-series overlap maximally (Oben, 2015), which allows the analyst to determine whether behaviors typically occur at the same time, with a given time lag, or unrelated from one another (see ; Xu et al., 2020 for overviews). Below, we provide a brief overview of the main methodological steps that are involved in a cross-recurrence quantification, but for a more comprehensive explanation, we refer to . First, to allow for CRQA, we sampled the existing gaze annotations into categorical time series at intervals of 100 ms. Concretely, for the start of each 100 ms timeslot, we indicated on which specific area of interest the participants were fixating their gaze. Next, recurrence is interpreted as two co-visitors either looking at the same artwork, text or label together or looking at each other at the same time. Importantly, two interlocutors looking at other potential areas of interest (e.g., walls, floors, other people) simultaneously is not considered recurrence. After sampling and categorizing the data in this way, we used the R-package developed by to perform the cross-recurrence analysis.
In line with existing research on gaze synchronization (e.g., ; Richardson and Dale, 2005), we also calculated a baseline level of synchronization to ensure that possible results are not driven by chance or external factors unrelated to the interaction. The latter is especially relevant within the context of a museum exhibition, as two people who are visiting the exhibition independently would inevitably also encounter the same artworks in a relatively similar order. While that would lead to gaze synchronization with a certain (arguably larger) time lag, that is not the kind of interactional behavior matching that we are interested in. To allow for a fair baseline comparison, we simulated a baseline synchronization level by reshuffling the time series of our data, or in other words randomly assigning every data point of the sampled dataset to a different position in time. In doing so, we generated 300 pairs of temporally randomized gaze annotations to perform separate CRQA analyses on. Given the particular spatial context of a museum exhibition as mentioned above, we specifically opted for a strict baseline calculation where we computed the baseline per AOI category (i.e., artworks, texts, labels or co-participants). As such, one person looking at a certain artwork and another looking at a different artwork is also labeled as synchronization, leading to a very strict benchmark. Since the average profile of the simulated CRQA analyses depicts the chance level of synchronization, the CRQA profile of the real dataset should be significantly higher than the averaged baseline plot to conclude that the synchronization is related to the interaction between the participants.
2.3 Qualitative method of analysis
Qualitatively, we use multimodal conversation analysis as our method of analysis. Rather than prioritizing verbal and paralinguistic resources, this perspective integrates the full spectrum of embodied resources available to participants. Specifically, the approach adopted here scrutinizes the interplay between these resources and how they are mobilized in a carefully timed way to contribute to the local negotiation of meaning (Mondada, 2013). This perspective thus refrains from presupposing fixed, pre-discursive functions for specific semiotic resources or treating them in isolation. Instead, we adopt a holistic approach in which the eye-tracking data can be integrated without prioritizing eye gaze over other bodily-visual resources. Importantly, to ensure the quality of these micro-oriented and qualitative analyses, we adhere to CA's fundamental principle of the next-turn proof procedure (Sacks et al., 1974). This procedure relies on the idea that “speakers display in their sequentially ‘next’ turns an understanding of what the ‘prior’ turn was about” (, p. 13). As such, this approach implicates that researchers do not aim to make claims about which meaning certain participants might have intended. Instead, the focus lies on scrutinizing how participants respond to each other and this emic perspective thus offers an empirical foundation for the analysis.
3 Analysis
To explore the methodological research objectives using the materials and methods explained above, the analysis proceeds in three steps. First, we conduct a quantitative data exploration using cross-recurrence quantification analysis to (i) determine whether joint attention occurs (above chance level) and, if so, on which areas of interest it does and (ii) explore the use of CRQA to gain insights into the establishment of joint attention in mobile interaction (§3.1). Second, a multimodal conversation analysis zooms in on the use of summonses (Stukenbrock and Balantani, 2025) to organize and negotiate joint attention (§3.2). Third, we provide an alternative perspective on the qualitative analysis by using a CRQA-inspired technique to visualize the patterns emerging from that micro-oriented analysis (§3.3).
3.1 Cross-recurrence quantification analysis
As a first step in the analysis, a quantitative data exploration aims to examine whether joint attention occurs and, if so, for which areas of interest it does. To this end, Figure 2 shows the recurrence rates from the CRQA analysis based on moments of interaction and averaging across all visitor pairs, either combining all areas of interest (top profile) or differentiating between the four categories at stake (i.e., artwork, text, label and co-participant) (lower profiles).
Figure 2
Combining all areas of interest, we observe a clear peak at t0, where the recurrence rate (RR, ranging from 0 to 1) increases to 0.16 compared to a baseline RR of 0.07. This indicates a distinct presence of simultaneous gaze synchronization, or in other words, joint attention between co-visitors across all annotated visitor pairs. To test whether the synchronization in the real interactions exceeds what would be expected by chance, we fitted a linear mixed-effects model with recurrence rate as the dependent variable, condition (real vs. baseline) as a fixed effect, and visitor pair as a random intercept. The model confirms that the prominent synchronization observed at t0 is statistically reliable: recurrence rates in the real data are significantly higher than in the shuffled baseline (t = 6.25, p < .001).
Zooming in on the individual AOI categories, the difference between the baselines indicates that there are more and/or longer stretches of gaze fixations on artworks and texts compared to labels and co-participants. In addition, Figure 2 shows a peak at t0 for ‘artwork’ (RR = 0.11 vs. baseline RR = 0.05), which is strongly supported by the mixed-effects model (t = 7.74, p < .001). For ‘text’, the RR at t0 reaches 0.04 compared to 0.01 at baseline. This difference is also significant, although with a smaller effect size (t = 3.12, p = .008). For both ‘label’ and ‘co-participant’, the curves for the real and baseline data closely overlap near zero. Concretely, both categories show a RR of approximately 0.01 at t0, with the mixed-effects model indicating small but statistically significant contrasts (label: t = 3.15, p = .004; co-participant: t = 3.28, p = .005). Overall, these findings thus suggest that there is a significant amount of gaze synchronization for all AOI categories. Finally, the pronounced lobes to the left and right of t0 (for the combined profile as well as for artworks and texts analyzed separately) indicate that, in most cases, one participant fixates on a given AOI before the other. As would be expected, this pattern suggests that moments of joint attention typically emerge when one individual is already looking at a given object and the second subsequently follows (which can be related to the gaze-cueing effect or the concept of gaze following/monitoring).
In line with interactionally oriented research on joint attention, this broader observation prompts a closer examination of how joint attention is organized in mobile interaction. To address this, a second exploratory analysis focuses specifically on the temporal dynamics of gaze shifts by considering only the onset of gaze fixations (which is operationalized as the first 500 ms of each fixation, following, for example, Oben et al., 2025). This approach allows the CRQA to more precisely capture potential leader-follower patterns in the organization of joint attention, whereas such effects may have been overshadowed in the previous analysis based on the full-duration annotations. Considering gaze onset only, a distinct peak next to t0 would indicate a systematic time lag or a pattern where one participant consistently follows the other with a specific reaction time. A peak at t₀ would instead suggest that gaze shifts by co-participants typically occur within 500 ms of each other. Lastly, multiple peaks would imply the occurrence of leader-follower behavior, but with varying reaction times and/or in two different directions. Given the unrestrained interactional set-up of the dataset, it is unlikely that a single, consistent ‘leader’ emerges within any visitor pair (as opposed to experimental designs with designated speakers and listeners, see for example Richardson and Dale, 2005). Consequently, to distinguish leader-follower dynamics in the establishment of joint attention in mobile interaction, we would expect to observe at least one peak left of t0 and one right of t0 on average, unless the recurrence of gaze shifts primarily occurs within a ±500 ms time window. To examine whether such patterns occur in the data, Figure 3 presents the recurrence rates of the second CRQA analysis, considering gaze onset only in a window of ±5 s, averaged across visitor pairs.
Figure 3
In line with Figure 2, the results in Figure 3 show a clear peak at t0 (RR = 0.009 compared to baseline RR = 0.003), which proves to be significant in the mixed-effects model (t = 8.42, p < 0.001). As such, the second CRQA analysis indicates that gaze shifts of co-visitors typically occur within 500 ms of each other. In addition, the plot illustrates two smaller peaks to the left and to the right of t0, possibly indicating leader-follower dynamics beyond this 500 ms window. However, the mixed-effects model reveals that these peaks (with time lags of approximately three seconds) do not illustrate a significant effect in terms of gaze shift synchronization (left peak: t = 1.82, p = 0.091; right peak: t = 1.47, p = 0.165). Nonetheless, a closer inspection of some of the individual CRQA plots per visitor pair illustrates that prominent peaks beyond t0 do occur. Consider, for example, the average plots of all interactions by dyads 58_59 and 39_40 in, respectively, panels (a) and (b) in Figure 4 below2.
Figure 4
For visitor pair 58_59 (a), the left side mainly shows a lobe extending from the zero peak, whereas the right side displays two prominent peaks at around +1.5 s and +3.5 s. For visitor pair 39_40 (b), the recurrence profile shows a peak on the right between approximately +1 s and +2.5 s, and a lobe emerging from the peak at t0 as well as a distinct peak ranging from roughly −1.5 s to −4 s on the left side of the plot. Overall, dyad 58_59 shows a more asymmetrical distribution, indicating a clearer tendency for one participant to lead as the rightside peaks are arguably more prominent than the left-side lobe extending from the zero peak. In contrast, dyad 39_40 exhibits a relatively symmetrical pattern with noticeable peaks on either side of the plot, suggesting that both participants alternate in leading and following each other's gaze direction. In addition, both CRQA profiles still include a clear peak at t0, which indicates that the average gaze pattern observed in Figure 3 (i.e., gaze shifts mainly occurring within 500 ms of each other) also frequently surfaces in the interactions of these visitor pairs.
However, the overall recurrence patterns clearly suggest that there is more involved than near-simultaneous gaze shifts and it remains unclear what exactly is happening interactionally that explains these peaks further away from t0. On the one hand, participants could actively be summoning their co-visitor's attention onto a certain AOI (cf. Stukenbrock and Balantani, 2025), using a variety of semiotic resources (including eye gaze, but also pointing gestures, demonstratives, directives, etc.), which could overshadow possible gaze shift synchronization effects. On the other hand, the patterns observed here might still be a residue of the spatial environment of the interaction, especially since the dataset is already filtered to only include gaze annotations overlapping with speech. In order to further examine this and analyze the negotiation of joint attention in mobile interaction from a more holistic perspective, the next section involves a micro-oriented qualitative analysis of two interactions by visitor pairs 39_40 and 58_59.
3.2 Multimodal conversation analysis
For the qualitative analysis, we selected two excerpts that involve clear instances of noticings and summonses used to initiate a shared focus of attention (Stukenbrock, 2020; Stukenbrock and Balantani, 2025), which can be related to a leader-follower dynamic in the interactional establishment of joint attention (cf. supra). Both excerpts take place in the same corner of the exhibition, which is schematized in Figure 5. It is important to point out that the portraits in artworks 1, 4, 5, 6 and 7 are of the same woman (i.e., the laundry woman of Peter Paul Rubens) and that this figure is also represented in artwork 3 (see 31). Similarly, artwork 2, which is a portrait of a man called Grapheus, also returns in artwork 3 (see 32).
Figure 5
In Excerpt 1, partners Lucy and Richard (i.e., dyad 58_59) are discussing some of the artworks portrayed in Figure 5. They have been looking at the first three paintings for a couple of minutes, when Richard opens up the interaction in line 1:

In line 1, Richard summons Lucy's attention by combining the use of a proper name with a perceptual directive (“look”) and by monitoring Lucy's gaze direction (notice Richard's gaze shift from aw2 to aw1, thereby aligning his gaze focus with Lucy's4). While Richard is thus requesting addressee gaze (; Stukenbrock, 2020), Lucy redirects her own gaze focus on aw31 instead of following Richard's gaze back to aw2 (l.1). She then latches onto Richard's summons with the news receipt marker “a:h ye:s” (), formulates a noticing (Stukenbrock, 2023) (“there she is”) and clarifies the deictic reference of “there” by pointing at aw31 (l.2). Richard does not engage with Lucy's noticing and instead shifts his gaze between aw1 and aw2, while latching onto Lucy's utterance and formulating an account for his summons (l.2-l.3). Concretely, he first looks and points at aw2 (Figure 6) while using a demonstrative “that” to refer to the figure in that painting and then shifts his gaze and pointing gesture towards aw32, combined with the deictic “here” (l.3). Richard is thereby establishing a link between the figures in aw2 and aw32 (“that woman5 you find here”, l.3) and summoning Lucy's attention onto this link through gaze shifts, pointing devices and deictics (Stukenbrock and Balantani, 2025). Lucy, who is still pointing at aw31 with her left hand, first uses her right hand to point at aw1 (Figure 6) while Richard verbally refers to aw2, before then following Richard's gaze direction towards aw2 (l.3). However, she then focuses her gaze on aw31 instead of aw32 (l.3) and negates Richard's noticing with a clear “no” (l.4). This negation is combined with a pointing gesture and a gaze shift towards aw1 and the use of a loudly pronounced demonstrative “this” to refer to the figure in aw1, providing an alternative mapping of two paintings (l.4). While Richard follows Lucy's gaze to aw1, he then looks back at aw32 (l.4), which leads to a negation of Lucy's utterance (l.5). So, Lucy and Richard have both formulated a noticing (the link between aw1 and aw31 and between aw2 and aw32, respectively) and tried to summon their partner's attention towards these noticings. But, while they both consider each other's ‘source paintings' (i.e., aw1 and aw2), they remain fixated on their own respective element within aw3 and thus fail to recognize the other's noticing.
Figure 6
In line 5, Richard, however, repeats his utterance of line 3 (including the same verbal construction and the same use of gaze and pointing gestures) and for the first time, Lucy redirects her gaze from aw31 to aw32. During a longer pause of four seconds, Lucy then shifts her gaze from aw32 to txt3 and Richard alternatingly looks at aw's 1, 2 and 3 (fixating on different elements of aw3, including aw32 but still excluding aw31) (l.6). Following the lack of response by Lucy, there is another attempt by Richard to summon his partner's attention onto his noticing (l.7). This time, Richard offers more information and clarifies that it is a figure with “gaze upwards”, pointing at aw2 while looking back and forth at aw2 and Lucy (l.7). During this utterance, Lucy gazes at aw2 and then aw32, thereby noticing Richard's attentional target, which is also verbally acknowledged with a news receipt marker “a:h” followed by three yeses (i.e., a display of shared perception, Stukenbrock and Balantani, 2025) (l.8).
Simultaneously, Richard is looking at aw3 again and notably focuses on various elements of the painting except aw31, whereas Lucy's gaze shifts back from aw32 to aw31 (l.8). During a pause of 1.1 s, Richard focuses his gaze on aw1 (l.9), which is followed by another attempt of Lucy's to summon his attention on her own initial noticing (l.10). In particular, she looks back at aw1, explicitly contrasts her noticing with Richard's (“but I see”) and then summons Richard's attention onto aw1 with an emphasized demonstrative and an embodied pointing device (Stukenbrock and Balantani, 2025) (l.10). Next, she shifts her gaze towards aw31, changes the direction of her pointing gesture and uses a deictic (“there”) to refer to the figure in aw31 (l.10). She ends her utterance with a confirmation request (“right” with rising intonation) and also re-acknowledges Richard's painting link by including the adverb “too” (l.10). Richard responds to this with a clarifying question (“where” with rising intonation in l.11), switching his gaze focus between aw1 and aw3, yet still not noticing aw31. In turn, Lucy (still pointing at aw31) now walks closer to aw3, responds with “there” and focuses her gaze on aw31 (l.12). For the first time, Richard's gaze now also fixates on aw31 (Figure 7) and he confirms this noticing with a loud “okay”, while quickly looking back at aw1. Richard, still going back and forth between aw1 and aw31, then utters a final assessment (“then it is good”) and thereby acknowledges this second joint focus of attention (l.14–16).
Figure 7
In conclusion, Excerpt 1 involves two competing foci of attention (Stukenbrock and Balantani, 2025), namely Richard's noticing of the link between aw2 and aw32 and Lucy's noticing of the relation between aw1 and aw31. While both interlocutors try to summon each other's attention on their own noticeable, using typical semiotic resources such as gaze cueing (McKay et al., 2021), perceptual directives, embodied pointing gestures, demonstratives and other deictics (Stukenbrock and Balantani, 2025), they initially both fail to establish a joint focus of attention. However, in the second half of the excerpt, the interactants repeat their summonses and eventually do succeed in cueing each other's visual attention on the right targets, in an embodied choreography involving gaze, gesture, movement and speech (Tulbert and Goodwin, 2011). In Section 3.1, we saw that the average CRQA plot for this visitor pair (Figure 4a) shows a rather asymmetric pattern in which Richard typically takes the lead (cf. the two distinct peaks on the right side of the plot). While Lucy's gaze also seems to cue Richard's (cf. the left-side lobe of the t0 peak), this happens to a lesser extent, as there are no distinct peaks on that side of the CRQA plot. In this fragment, we can observe a similar dynamic where Richard's noticeable takes the upper hand compared to Lucy as he only considers his co-visitor's noticing after repeating his own summonses until there is a joint focus of attention on his link between the artworks.
In the next excerpt, two partners (Steven and Eleanor, dyad 39_40) are discussing the same paintings as the artworks from the previous analysis. Before the start of the fragment, Eleanor had already attended this corner of the exhibition alone, while listening to the audioguide and reading some of the information texts. When Steven joins this part of the exhibition room, he shifts his gaze between artworks 1 and 3 (as well as their labels) for approximately half a minute, before Eleanor walks towards him and starts to retell “the story” behind the paintings (line 1):

In the first line, Eleanor is looking at aw1 when she starts the interaction with a topic-opener “so”, upon which Steven shifts his gaze towards her. While pointing and looking at the first artwork, Eleanor continues her multi-unit turn by stating that “that is the laundry woman” with a slightly rising intonation (l.2). During this utterance, Steven's gaze briefly pauses on aw3 and txt3 before he also looks at aw1 and utters a continuer (“mhm”) (l.3). Thus, Eleanor has summoned Steven's attention onto the first painting by combining a gaze shift with a pointing gesture and the demonstrative that (Stukenbrock and Balantani, 2025). Subsequently, she shifts her gaze to the copy of aw1 in aw3 (aw31) (l.3) and finishes her sentence by clarifying that it is the laundry woman “of Peter Paul Rubens” (l.4). Although Eleanor is still pointing at aw1, Steven already follows her gaze in the direction of aw3 and, notably, briefly fixates his gaze on aw31 in particular. In terms of gaze behavior, there is thus already a moment of joint attention on aw31 (Figure 8), while the verbal and the gestural level of the interaction is still concerned with aw1. Next, Eleanor starts to look at aw4 and there is a brief pause of 0.6 s (l.5), after which Steven says “okay” and redirects his gaze to a wall on the left side of aw4-aw7 (l.6). Latching onto this continuer/confirmation, Eleanor now relates the laundry woman from aw1 to aw4 and then aw3 (cf. “they are also there and there” in l.7). During these two deictic “there” references, Eleanor both looks and points at the paintings she is referring to (l.7). Interestingly, Steven has already directed his gaze at aw6, aw7 and then back at aw3, before Eleanor's “there” deictics and pointing gestures (l.7). Considering that aw6 and aw7 are similar to aw4 (as they are all small paintings of Rubens' laundry woman on the left part of this exhibition corner), it is interesting that here too, Steven had already followed the (broad) direction of Eleanor's gaze before the more explicit verbal and gestural resources used to redirect his attention. The same applies to aw3, which leads to a joint focus of attention in the short pause before Eleanor's second “there” reference and pointing gesture in line 7. During the following pause of 0.8 s, Eleanor shifts her gaze to her co-participant (l.8), who responds with another continuer (“mhm”) in line 9. Eleanor then laughingly concludes that the artists were “copy pasting” and shifts her gaze between aw1 and her partner (l.10), which can be related to previous research indicating an increase in gaze monitoring or gaze shifts towards addressees at the end of jocular comments (see e.g., ). Simultaneously, Steven shifts his gaze from aw3 to txt3 and then back to aw3 (l.10). In line 11, he then utters the news receipt marker “a:h ye:s” and looks back at Eleanor (l.11).
Figure 8
In overlap, Eleanor starts another account for her joke (“because”) and starts to look and point at aw2 (i.e., the only portrait of someone other than Rubens' laundry woman), while claiming that “that is also Grapheus” (l.12). Thus, she uses another combination of a pointing gesture and a demonstrative to request addressee gaze (Stukenbrock and Balantani, 2025), which leads to a moment of joint attention as Steven shifts his gaze towards aw2 as well6 (l.12, Figure 9). Afterwards, there is a pause of 0.9 s during which Eleanor shifts her gaze towards aw3 and starts to point in the direction of aw32 (i.e., the copy of aw2 in aw3) (l.13). While Steven directs his gaze at aw4, 5 and 6 instead, Eleanor starts to continue her utterance about Grapheus (“and they are there then”) (l.14), but this is interrupted as Eleanor's pointing gesture almost reaches another visitor's face. After briefly looking at this visitor and apologizing (l.15), Eleanor repeats herself and finishes her utterance in line 16 (“they are there then also in”), clarifying the deictic reference of “there” with a pointing gesture towards aw32. Steven already shifted his gaze at aw32 at the end of Eleanor's apology and is now looking at other elements of aw3, when Eleanor looks back at aw3 as well (l.15–l.16). In overlap with this, Steven utters another continuer (“mhm” in l.17), after which he redirects his gaze at txt3 again. At the end of the excerpt, there is another brief pause, after which Steven says “okay” and shifts his gaze to aw2 (l.19). Latching onto this, Eleanor concludes her recount of “the story” (l.1) by saying “that's what they said” (while looking at the main information text, txt3) and the topic is closed off.
Figure 9
In conclusion, Excerpt 2 involves two different patterns in the organization of a joint focus of attention between the interlocutors. On the one hand, we observed the typical combination of gaze cueing, embodied pointing and demonstratives or deictics (Stukenbrock and Balantani, 2025) to establish moments of joint attention. On the other hand, there were also multiple instances where gaze cueing alone sufficed, and yet the gestural and the verbal tier still occurred. A possible account for this is that Eleanor formulates her (multi-unit turn) summonses rather slowly (notice the pauses in lines 1, 2, 5, 7, 8, 10 and 13 as well as the continuers by Steven in lines 3, 6 and 9), which leaves a lot of room for the multimodal layer of the interaction (and especially Eleanor's gaze cueing) to take the upper hand. Another explanation might also be the strong coupling of deictics and aligned gestures regarding demonstrative reference, which can form multimodal packages or Gestalts (Stukenbrock, 2020). Considering the short duration of the moments of joint attention as well as the lack of engagement beyond minimal responses by Steven, it also makes sense that Eleanor completes these recipient-designed Gestalts. In general, the addressee displays markedly less engagement and commitment in this fragment compared to the previous one. While the first excerpt illustrated a case of two competing foci of attention in which there is more at stake for both of the interlocutors, this excerpt involves one participant who is sharing information or expressing noticings and another who is mostly a passive recipient of that information. In comparison to the average CRQA plot for this visitor pair (Figure 4b), this fragment thus illustrates a rather asymmetric interaction in which Eleanor consistently summons Steven's attention, whereas the CRQA plot shows a relatively symmetric recurrence pattern in which both of them seem to cue each other's gaze direction.
In general, the qualitative analysis illustrated a build-up from multimodal summons packages (including gaze cueing, pointing gestures and verbal deictics) failing (Excerpt 1), succeeding (Excerpts 1, 2) and not being fully required (Excerpt 2) in the establishment of joint attention. Furthermore, in terms of leader-followership in the organization of a joint focus of attention (i.e., Who is summoning whom?), the first excerpt includes an example of two co-visitors alternatingly requesting each other's visual attention and also alternatingly responding to that. In comparison, the second excerpt shows an example of one participant taking the lead, while the other only formulates continuers and minimal responses. Compared to the average CRQA plots as observed in Figure 4 (§3.1), the first excerpt thus includes an example of a fragment that is in line with the average gaze synchronization pattern across the museum visit, whereas the second excerpt illustrates an example of a fragment that deviates from the average pattern in the conversation. Therefore, in the next analytical section, we aim to explore how a close-recurrence analysis for these specific fragments alone (instead of an average CRQA based on all interactions during the museum visit such as in Figure 4) could deepen our understanding of the interactional dynamics at play here.
3.3 Close-recurrence analysis
In this section, we propose the idea of a close-recurrence analysis as a visualization technique that is based on the underlying principles of CRQA. Essentially, we conduct a small-scale CRQA of the gaze behaviors from the qualitative analysis to provide an additional perspective on that micro-oriented analysis. For this close-recurrence analysis, we selected the fragments as portrayed in Section 3.2 (i.e., from the beginning of the first utterance until the end of the last utterance), which translates to 22 s for the first excerpt and 27 s for the second excerpt. In ELAN, we differentiated between aw31 and aw32 in the gaze fixations, thereby refining the annotations in comparison to the CRQA analyses from Section 3.1. As opposed to the average CRQA analyses, the baselines were not calculated per category, given the more specific interest in gaze synchronization on particular AOI's here. Finally, all of the close-recurrence analyses are based on the onsets of the gaze fixations (i.e., the first 500 ms, cf. Section 3.1) and a window size of 2.5 s (to fully focus on interactional coordination rather than incidental leader-followership related to the spatial environment).
As a result, Figure 10 focuses on the gaze behavior of visitor pairs 58_59 (Excerpt 1) and 39_40 (Excerpt 2) in panels (a) and (b) respectively. In line with the findings from the qualitative analysis, the plot in panel (a) shows that both Richard and Lucy follow each other's gaze directions. The left side of the plot includes one peak indicating Richard following Lucy's gaze fixations with a reaction time of about 0.5–1.5 s. On the right, there are two peaks reflecting Lucy's following of Richard's gaze, including a smaller peak within one second after t0 and a prominent peak between one to two seconds after t0. While it is clear that they both orient to each other's visual focus of attention, Lucy seems to follow Richard more than the other way around (given the higher recurrence rate of the right-side peak), which is also supported by the qualitative analysis (cf. Excerpt 1). Also in accordance with the observations from the qualitative analysis, panel (b) shows a prominent peak of Steven's gaze fixations following Eleanor's within one second of Eleanor's gaze shifts towards a particular AOI. So, in terms of gaze behavior, there is indeed one designated leader and one follower within the timespan of the excerpt. Interestingly, these CRQA plots thus also show that Steven's reaction time is quicker than Richard's and Lucy's, which arguably relates to the observation that gaze-cueing alone often sufficed to request Steven's attention (cf. Excerpt 2). Another possible explanation for the difference in reaction times between the two plots is that there is more ‘at stake’ for Richard and Lucy compared to Steven (see above), which might cause a delay in terms of response times as Richard and Lucy are navigating their gaze fixations between two competing foci of attention.
Figure 10
4 Discussion
This study emphasizes the benefits of a mixed-methods approach for investigating joint attention in general and in mobile interaction in particular. The quantitative data exploration using CRQA provided a broad overview of where and when gaze synchronization and joint attention occurred in our dataset of joint museum visits. As expected, applying CRQA to naturally occurring mobile interactions also proved to be less straightforward than in the controlled, task-based settings in which the method is typically employed. In particular, it is difficult to observe stable and easily interpretable recurrence profiles when considering the organization of joint attention averaged across all visitor pairs. This may be due to a variety of factors related to mobile interactions in general or joint museum visits in particular: individual variation between the participants, interpersonal factors such as the mutual relation between the co-visitors (e.g., friends vs. partners), contextual factors such as the crowdedness of the exhibition rooms or spatial factors including the complexity of mobile interactions in a dynamic environment with many visual input sources. Therefore, micro-oriented qualitative analyses are especially useful here to further interpret the results from a general CRQA analysis. On the one hand, the qualitative analysis is able to zoom in on the findings from the quantitative data exploration (e.g., Excerpt 1). On the other hand, it is also interesting to explore fragments that deviate from their corresponding average CRQA plot (e.g., Excerpt 2) to provide nuance to the quantitative analysis.
Importantly, this study also demonstrates that quantitative methods such as CRQA are not only useful for large datasets or wide-reaching analyses. While the analytical workflow that includes a general quantitative analysis and an ensuing micro-oriented qualitative analysis is already an established practice in mixed-methods research, we show how an adaptation of a quantitative method can also be relevantly applied to a qualitative analysis. In particular, we have shown how a close-recurrence analysis, based on the underlying principles of CRQA, can function as a visualization technique to clarify as well as substantiate qualitative results. As such, a version of CRQA can still enhance our understanding of the interactional dynamics that are at play in the negotiation of joint attention in mobile interaction. In conclusion, while CRQA works well to obtain general results in experimental settings, it struggles to provide clearly interpretable results in a more free and mobile interactional setting. However, if we reduce the scope of the input data significantly, an alternative version of CRQA proves to be useful in the study of gaze in naturally occurring interactions. Nevertheless, a limitation of our close-recurrence approach is that implementing CRQA (even on a small scale) requires substantial methodological preparation and computational resources, despite being used here primarily as a visualization tool. Developing a more accessible tool for conducting close-recurrence analyses would therefore be a valuable direction for future research.
More broadly, we have also started to explore the use of close-recurrence analyses to examine the temporal relations between gaze and other resources emerging from the qualitative analysis (such as pointing gestures and verbal deictics), which has indicated a promising avenue for further research. In line with this, future studies could also examine the temporal relations between multiple semiotic resources on a larger scale and explore multidimensional recurrence techniques capable of visualizing several semiotic resources simultaneously. In addition, it would be interesting to explore the use of close-recurrence analysis as an automatic pointer towards specific moments where leader-follower patterns appear to be strong, or in other words, employ a quantitative step to easily identify potentially interesting sequences to further analyze in a qualitative analysis. Such work would further refine our understanding of how joint attention is organized in naturalistic environments and benefit the methodological integration of quantitative and qualitative approaches in the study of gaze in interaction.
Statements
Data availability statement
The datasets presented in this article are not readily available because, in compliance with the data collection procedure that was approved by the KU Leuven Social and Societal Ethics Committee (file nr. G-2023-7074), it is not possible to publish the original eye-tracking data and videorecordings of the dataset. However, the gaze annotations used to conduct the cross-recurrence analyses in this study, as well as additional CRQA plots for the individual visitor pairs can be made available upon request by interested researchers. Requests to access the datasets should be directed to julie.janssens1@kuleuven.be.
Ethics statement
The studies involving humans were approved by KU Leuven Social and Societal Ethics Committee (file nr. G-2023-7074). The studies were conducted in accordance with the local legislation and institutional requirements. The participants provided their written informed consent to participate in this study. Written informed consent was obtained from the individual(s) for the publication of any potentially identifiable images or data included in this article.
Author contributions
JJ: Conceptualization, Data curation, Formal analysis, Funding acquisition, Investigation, Methodology, Visualization, Writing – original draft. BO: Conceptualization, Methodology, Supervision, Writing – review & editing. DV: Conceptualization, Methodology, Supervision, Writing – review & editing. GB: Conceptualization, Methodology, Project administration, Supervision, Writing – review & editing.
Funding
The author(s) declared that financial support was received for this work and/or its publication. This research was supported by the Research Foundation Flanders—grant number 1147626N.
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The author(s) declared that generative AI was used in the creation of this manuscript. Generative AI (CoPilot) was used for text revision of a limited amount of paragraphs.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Footnotes
1.^In terms of interpersonal relations, these 15 visitor pairs include eight pairs of romantic partners, two sets of colleagues, one pair of friends and four dyads with a family relation (i.e. one pair of sisters-in-law, one pair of cousins, a mother-daughter duo and a father-daughter duo). Because of this unbalanced distribution, we have not tested whether the interpersonal relation of the visitor pairs has an effect on their gaze behavior. For the sake of completeness, we do mention that dyads 39_40 and 58_59 (which will be analyzed in detail) both consist of romantic partners.
2.^To select the visitor pairs for this example, we first selected all average CRQA plots with a baseline of at least 0.002 to ensure a sufficient level of synchronization. This was the case in eight out of the fifteen visitor pairs that were analyzed. Among these eight dyads, three visitor pairs showed noticeably prominent peaks beyond t0. From those three, we then selected visitor pairs 39_40 and 58_59 because they illustrate a symmetric versus an asymmetric recurrence pattern.
3.^The abbreviations in the multimodal transcription lines represent the different types of AOI: aw=artwork, lbl=label, txt=text, part=co-participant. The numbers refer to the different artworks as illustrated in Figure 5 and, for example, txt1 thus represents an information text that accompanies aw1.
4.^Note that Richard's gaze shift towards aw1 occurs almost immediately after Lucy had started to look at that painting (l.1). This swift organization of gaze synchronization thus illustrates the average pattern observed in Figure 3, which denoted that gaze shifts resulting in a moment of joint attention typically take place within a short timeframe of 500 ms.
5.^Note that Richard is wrongfully referring to Grapheus as a woman here.
6.^Note that this is another illustration of the average gaze pattern observed in Figure 3, as Steven's gaze shift towards aw2 occurs within a short timespan of Eleanor's gaze shift towards that painting (l.12).
References
1
AuerP. (2021). Turn-allocation and gaze: a multimodal revision of the “current-speaker-selects-next” rule of the turn-taking system of conversation analysis. Discourse. Stud.23 (2), 117–140. 10.1177/1461445620966922
2
AuerP.LanerB. (2025). “Laughter and gaze among talkers on a walk,” in Pragmatics & Beyond New Series, Vol. 351, eds. ZimaE.StukenbrockA. (Amsterdam/Philadelphia: John Benjamins), 208–242. 10.1075/pbns.351.08aue
3
AuerP.LanerB.PfeifferM.BotschK. (2024). “Noticing and assessing nature: a multimodal investigation of the format “perception imperative+exclamative” based on mobile eye-tracking data,” in New Perspectives in Interactional Linguistic Research, eds. SeltingM.Barth-WeingartenD. (Amsterdam/Philadelphia: John Benjamins), 245–275. 10.1075/slsi.36
4
AuerP.ZimaE. (2021). On word searches gaze and co-participation. Gesprächsforsch. Online-Z. Verb. Interakt.22, 390–425.
5
BalantaniA. (2021). Reference construction in interaction: the case of type-indicative “so”. J. Pragmat.181, 241–258. 10.1016/j.pragma.2021.05.024
6
BalantaniA.LázaroS. (2021). Joint attention and reference construction: the role of pointing and “so”. Lang. Commun.79, 33–52. 10.1016/j.langcom.2021.04.002
7
BotschK.AuerP.LanerB.PfeifferM. (2025). “Joint attention without language? On intersubjectivity and the joint experience of nature,” in Mobile Eye Tracking: New Avenues for the Study of Gaze in Social Interaction, eds. ZimaE.StukenbrockA. (Amsterdam/Philadelphia: John Benjamins), 277–310. 10.1075/pbns.351
8
BrôneG.ObenB. (2015). Insight interaction: a multimodal and multifocal dialogue corpus. Lang. Resour. Eval.49 (1), 195–214. 10.1007/s10579-014-9283-2
9
BrôneG.ObenB. (Eds.) (2018). Eye-tracking in Interaction: Studies on the Role of Eye Gaze in Dialogue, (Amsterdam/Philadelphia: John Benjamins). 10.1075/ais.10
10
BrôneG.ObenB. (2023). “Mobile eye-tracking FOR multimodal interaction analysis,” in The Routledge Handbook of Experimental Linguistics, eds. GygaxP.ZuffereyS. (London: Routledge), 283–298. 10.4324/9781003392972-21
11
CaruanaN.Stieglitz HamH.BrockJ.WoolgarA.KlothN.PalermoR.et al (2018). Joint attention difficulties in autistic adults: an interactive eye-tracking study. Autism22 (4), 502–512. 10.1177/1362361316676204
12
ClarkA. T.GergleD. (2011). “Mobile dual eye-tracking methods: challenges and opportunities,” in Proceedings of international workshop on dual eye-tracking, 7.
13
CocoM. I.DaleR. (2014). Cross-recurrence quantification analysis of categorical and continuous time series: an R package. Front. Psychol.5, 510. 10.3389/fpsyg.2014.00510
14
DaleR.KirkhamN. Z.RichardsonD. C. (2011). “How two people become a tangram recognition system,” in Proceedings of the European conference on computer-supported cooperative work.
15
de VriesC.ObenB.BrôneG. (2021). Exploring the role of the body in communicating ironic stance. Lang. Modalities1, 65–80. 10.3897/lamo.1.68876
16
de VriesC.ObenB.BrôneG. (2023). On target. On the role of eye-gaze during teases in face-to-face multiparty interaction. Interact. Humor10, 53–86. 10.1515/9783110983128-003
17
DiesselH. (2006). Demonstratives, joint attention, and the emergence of grammar. Cogn. Linguist.17 (4), 463–489. 10.1515/COG.2006.015
18
FusaroliR.KonvalinkaI.WallotS. (2014). “Analyzing social interactions: the promises and challenges of using cross recurrence quantification analysis,” in Translational Recurrences, Vol. 103, eds. MarwanN.RileyM.GiulianiA.WebberC. L. (Cham/Heidelberg/New York/Dordrecht/London: Springer), 137–155. 10.1007/978-3-319-09531-8_9
19
GoodwinC. (1981). Conversational Organization: Interaction Between Speakers and Hearers. London: Academic Press.
20
GoodwinM. H.GoodwinC. (2012). Car talk: integrating texts, bodies, and changing landscapes. Semiotica.2012 (191), 257–286. 10.1515/sem-2012-0063
21
HadelichK.CrockerM. W. (2006). “Gaze alignment of interlocutors in conversational dialogues,” in Proceedings of the 2006 symposium on eye tracking research & applications, ETRA ‘06, 38. 10.1145/1117309.1117322
22
HeritageJ. (1985). “A change-of-state token and aspects of its sequential placement,” in Structures of Social Action, ed. Maxwell AtkinsonJ. (Cambridge: Cambridge University Press), 299–345. 10.1017/CBO9780511665868.020
23
HutchbyI.WooffittR. (Eds.) (2008). Conversation Analysis, 2nd Edn, (Cambridge: Polity Press).
24
JeffersonG. (2004). “Glossary of transcript symbols with an introduction,” in Conversation Analysis; Studies from the First Generation, ed. LernerG. H. (Philadelphia: John Benjamins), 13–31.
25
KendonA. (1967). Some functions of gaze-direction in social interaction. Acta Psychol. (Amst)26, 22–63. 10.1016/0001-6918(67)90005-4
26
KendrickK. H.HollerJ. (2017). Gaze direction signals response preference in conversation. Res. Lang. Soc. Interact.50 (1), 12–32. 10.1080/08351813.2017.1262120
27
KesselheimK. W.BrandenbergerC.HottigerC. (2021). How to notice a tsunami in a water tank: joint discoveries in a science center. Gesprächsforsch. Online-Z. Verb. Interakt.22, 87–113.
28
LandisJ. R.KochG. G. (1977). The measurement of observer agreement for categorical data. Biometrics33, 159–174. 10.2307/2529310
29
LanerB. (2022). Guck mal der Baum—zur Verwendung von Wahrnehmungsimperativen mit und ohne ‘mal’. Gesprächsforsch. Online-Z. Verb. Interakt.23, 1–35.
30
LouwerseM. M.DaleR.BardE. G.JeuniauxP. (2012). Behavior matching in multimodal communication is synchronized. Cogn. Sci.36 (8), 1404–1426. 10.1111/j.1551-6709.2012.01269.x
31
McKayK. T.GraingerS. A.CoundourisS. P.SkorichD. P.PhillipsL. H.HenryJ. D. (2021). Visual attentional orienting by eye gaze: a meta-analytic review of the gaze-cueing effect. Psychol. Bull.147 (12), 1269–1289. 10.1037/bul0000353
32
MondadaL. (2013). Embodied and spatial resources for turn-taking in institutional multi-party interactions: participatory democracy debates. J. Pragmat.: Convers. Anal. Stud. Multimodal Interact.46 (1), 39–68. 10.1016/j.pragma.2012.03.010
33
MondadaL. (2018). Multiple temporalities of language and body in interaction: challenges for transcribing multimodality. Res. Lang. Soc. Interact.51 (1), 85–106. 10.1080/08351813.2018.1413878
34
NeiderM. B.ChenX.DickinsonC. A.BrennanS. E.ZelinskyG. J. (2010). Coordinating spatial referencing using shared gaze. Psychon. Bull. Rev.17 (5), 718–724. 10.3758/PBR.17.5.718
35
ObenB. (2015). Modelling interactive alignment: a multimodal and temporal account. (Doctoral dissertation). Leuven:University of Leuven.
36
ObenB.de VriesC.BrôneG. (2025). “Mobile eye-tracking and mixed-methods approaches to interaction analysis,” in Mobile Eye Tracking: New Avenues for the Study of Gaze in Social Interaction, eds. ZimaE.StukenbrockA. (Amsterdam/Philadelphia: John Benjamins), 100–127. 10.1075/pbns.351
37
RichardsonD. C.DaleR. (2005). Looking to understand: the coupling between speakers’ and listeners’ eye movements and its relationship to discourse comprehension. Cogn. Sci.29 (6), 1045–1060. 10.1207/s15516709cog0000_29
38
RichardsonD. C.DaleR.KirkhamN. Z. (2007). The art of conversation is coordination. Psychol. Sci.18 (5), 407–413. 10.1111/j.1467-9280.2007.01914.x
39
RichardsonD. C.DaleR.TomlinsonJ. M. (2009). Conversation, gaze coordination, and beliefs about visual context. Cogn. Sci.33 (8), 1468–1482. 10.1111/j.1551-6709.2009.01057.x
40
SacksH.SchegloffE. A.JeffersonG. (1974). A simplest systematics for the organization of turn-taking for conversation. Language. (Baltim)50 (4), 696–735. 10.1353/lan.1974.0010
41
ShiL.SticklerU. (2021). Eyetracking a meeting of minds: teachers’ and students’ joint attention during synchronous online language tutorials. J. China Comput.-Assist. Lang. Learn1 (1), 145–169. 10.1515/jccall-2021-2006
42
StukenbrockA. (2018a). “Forward-looking: where do we go with multimodal projections?,” in Time in Embodied Interaction: Synchronicity and Sequentiality of Multimodal Resources, eds. DeppermannA.StreeckJ. (Amsterdam: John Benjamins), 31–68. 10.1075/pbns.293.01stu
43
StukenbrockA. (2018b). “Mobile dual eye-tracking in face-to-face interaction: the case of deixis and joint attention,” in Eye-tracking in Interaction: Studies on the Role of Eye Gaze in Dialogue, eds. BrôneG.ObenB. (Amsterdam/Philadelphia: John Benjamins), 265–300. 10.1075/ais.10.11stu
44
StukenbrockA. (2020). Deixis, meta-perceptive gaze practices, and the interactional achievement of joint attention. Front. Psychol.11, 1779. 10.3389/fpsyg.2020.01779
45
StukenbrockA. (2023). Temporality and the cooperative infrastructure of human communication: noticings to delay and to accelerate onward movement in mobile interaction. Lang. Commun.92, 33–54. 10.1016/j.langcom.2023.06.003
46
StukenbrockA.BalantaniA. (2025). “When the establishment of joint attention becomes problematic: how participants manage divergent and competing foci of attention,” in Mobile Eye Tracking: New Avenues for the Study of Gaze in Social Interaction, eds. ZimaE.StukenbrockA. (Amsterdam/Philadelphia: John Benjamins), 243–276. 10.1075/pbns.351
47
StukenbrockA.DaoA. N. (2019). “Joint attention in passing: what dual Mobile eye tracking reveals about gaze in coordinating embodied activities at a market,” in Embodied Activities in Face-to-face and Mediated Settings: Social Encounters in Time and Space, eds. ReberE.GerhardtC. (Cham: Springer), 177–213. 10.1007/978-3-319-97325-8_6
48
TomaselloM.CarpenterM. (2007). Shared intentionality. Dev. Sci.10 (1), 121–125. 10.1111/j.1467-7687.2007.00573.x
49
TulbertE.GoodwinM. H. (2011). “Choreographies of attention: multimodality in a routine family activity,” in Embodied Interaction: Language and Body in the Material World, eds. StreeckJ.GoodwinC.LeBaronC. (New York: Cambridge University Press), 79–92. ISBN: 978-0521895637.
50
VertegaalR.SlagterR.Van Der VeerG.NijholtA. (2001). “Eye gaze patterns in conversations: there is more to conversational agents than meets the eyes,” in Proceedings of the SIGCHI conference on human factors in computing systems, 301–308. 10.1145/365024.365119
51
Vom LehnD. (2013). Withdrawing from exhibits: the interactional organisation of museum visits. Interact. Mobil.: Lang. Body Motion20, 65–90. 10.1515/9783110291278.65
52
VrzakovaH.AmonM. J.StewartA. E. B.D'MelloS. K. (2019). “Dynamics of visual attention in multiparty collaborative problem solving using multidimensional recurrence quantification analysis,” in Proceedings of the 2019 CHI conference on human factors in computing systems, 1–14. 10.1145/3290605.3300572
53
WittenburgP.BrugmanH.RusselA.KlassmannA.SloetjesH. (2006). “ELAN: a professional framework for multimodality research,” in Proceedings of the fifth international conference on language resources and evaluation, LREC ‘06, 1556–1559.
54
XuT. L.de BarbaroK.AbneyD. H.CoxR. F. A. (2020). Finding structure in time: visualizing and analyzing behavioral time series. Front. Psychol.11, 1457. 10.3389/fpsyg.2020.01457
55
ZimaE.AuerP.RühlemannC. (2025). “Why research on gaze in social interaction needs mobile eye tracking,” in Mobile Eye Tracking: New Avenues for the Study of Gaze in Social Interaction, eds. ZimaE.StukenbrockA.Amsterdam/Philadelphia: (John Benjamins), 24–66. 10.1075/pbns.351.02zim
56
ZimaE.StukenbrockA. (Eds.) (2025). Mobile Eye Tracking: New Avenues for the Study of Gaze in Social Interaction, (Amsterdam/Philadelphia: John Benjamins). 10.1075/pbns.351
57
ZimaE.WeißC.BrôneG. (2019). Gaze and overlap resolution in triadic interactions. J. Pragmat.140, 49–69. 10.1016/j.pragma.2018.11.019
Summary
Keywords
close-recurrence analysis, cross-recurrence quantification analysis, joint attention, mixed-methods, mobile eye-tracking, mobile interaction, multimodal conversation analysis
Citation
Janssens J, Oben B, Van De Mieroop D and Brône G (2026) Close-recurrence analysis: a mixed-methods tool for studying joint attention in mobile interaction. Front. Commun. 11:1821332. doi: 10.3389/fcomm.2026.1821332
Received
02 March 2026
Revised
04 May 2026
Accepted
12 May 2026
Published
29 May 2026
Volume
11 - 2026
Edited by
Antonio Bova, Catholic University of the Sacred Heart, Italy
Reviewed by
Ottilie Tilston, Université de Lausanne, Switzerland
Maíra Avelar, Universidade Estadual do Sudoeste da Bahia, Brazil
Updates
Copyright
© 2026 Janssens, Oben, Van De Mieroop and Brône.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Julie Janssens julie.janssens1@kuleuven.be
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.