Abstract
Referring expression generation (REG) presents the converse problem to visual search: given a scene and a specified target, how does one generate a description which would allow somebody else to quickly and accurately locate the target?Previous work in psycholinguistics and natural language processing has failed to find an important and integrated role for vision in this task. That previous work, which relies largely on simple scenes, tends to treat vision as a pre-process for extracting feature categories that are relevant to disambiguation. However, the visual search literature suggests that some descriptions are better than others at enabling listeners to search efficiently within complex stimuli. This paper presents a study testing whether participants are sensitive to visual features that allow them to compose such “good” descriptions. Our results show that visual properties (salience, clutter, area, and distance) influence REG for targets embedded in images from the Where's Wally? books. Referring expressions for large targets are shorter than those for smaller targets, and expressions about targets in highly cluttered scenes use more words. We also find that participants are more likely to mention non-target landmarks that are large, salient, and in close proximity to the target. These findings identify a key role for visual salience in language production decisions and highlight the importance of scene complexity for REG.
Introduction
Cognitive science research in the domains of vision and language faces similar challenges for modeling the way people use and integrate information. For modeling people's interpretation of visual scenes and for accounting for their linguistic descriptions of such scenes, both fields must address the ways that local cues are integrated with larger contextual cues and the ways that different tasks guide people's strategies.
Despite these seemingly interlinked problem domains, vision and language have largely been studied as separate fields. Where intersections do occur, there is evidence that the way viewers make sense of a visual scene does indeed guide the language they use to describe it – visual information influences which objects speakers identify as important enough to mention and how they characterize the relationships between those objects (Coco and Keller, ; Clarke et al., submitted). Likewise, language itself acts as a strong gaze cue – listeners' eye movements in psycholinguistic eye-tracking experiments reflect their real-time language comprehension (Tanenhaus et al., ). Existing studies at the vision ~ language interface have succeeded in incorporating complex visual stimuli or complex linguistic tasks, but rarely both, and the conclusions from that previous work have assigned a limited role to vision in language production. This paper considers the question of how the language people produce in a complex referential task is influenced by the properties of a complex visual scene. Specifically, participants in our study were asked to describe individuals in illustrated crowd scenes; we then test whether the elicited descriptions reflect the visual properties of the targets themselves and of the complex scenes in which those targets appear.
In order to generate a natural and contextually appropriate description of a target object, a speaker must identify what properties of that object are relevant in context and what kinds of descriptions would help a listener identify that object. Understanding what people do in such tasks provides clues for improving natural language processing (NLP) systems which generate such descriptions automatically (Viethen and Dale, ; Krahmer and van Deemter, ). This task, in which a person or NLP system builds a linguistic expression to pick out a particular object in context, fits under the interdisciplinary (psycholinguistics and NLP) domain of Referring Expression Generation (REG). In order to create an appropriate description, the viewer must gather perceptual information and then compose an expression that adheres to a set of linguistic constraints. This can be thought of as the converse problem to visual search, in which an observer is given a description of the target and then has to locate it within a visual scene.
As will be expanded on in the background sections on REG and visual perception, previous work has primarily focused on models of linguistic complexity or visual complexity but not both. Vision studies have kept the language task simple (“Describe what you see” or even just “Look at this scene”) and analyze effects from factors such as visual salience and display time (see Figure 1; and also Fei-Fei et al., ). On the other hand, REG studies have kept visual stimuli simple by using a small number of objects or a restricted number of feature dimensions while analyzing a more complex task (“Describe the highlighted object such that a listener could figure out which object you intended”), with the goal of evaluating the visual properties that people mention in distinguishing objects from one another (see Figure 2).
Figure 1
Figure 2
An open question is how the conclusions from such studies will scale. Given recent REG work which has concluded that the role for visual properties in complex REG tasks is small (Beun and Cremers,
In our study, we investigate the role of perception in REG using images from the children's book Where's Wally, published in the US as Where's Waldo. These images are an order of magnitude more complex than the arrays of geometric objects typically used in referring expression and visual search studies, with images containing many dozens of objects and people. In our task, viewers produce a description of one highlighted target in each scene. We demonstrate an important role for visual salience (Toet,
Referring expression generation
Early work in REG focused on the balance between brevity and descriptive adequacy – how to construct minimalist expressions that uniquely pick out the intended referent (Dale and Reiter,
The Incremental Algorithm and VOA fit within a larger psy-cholinguistics literature on audience design and common ground (Clark and Wilkes-Gibbs,
Given that speakers overspecify, REG algorithms are faced with the questions of what a natural-sounding referring expression should contain and how long the expression should be. For example, do speakers focus on the target object itself or do they recruit other objects in relational descriptions in order to convey how to find the target in the larger picture? As we will show in our study, as the salience of the target object decreases, participants include more descriptions of other objects which they use as landmarks. For example, although the character Wally in each Where's Wally image is unique, an unambiguous description of him (“the guy with the red hat and the black-and-white striped shirt”) is insufficient to help a listener pick out the intended referent. Rather, what participants rely on are visual properties of the scene that can be referenced in the description of how to find the intended referent. The next section reviews research that has explicitly asked how visual properties influence speakers' referring expressions, and we point to the limitations with using small-scale visual scenes to answer that question.
Influence of visual properties in REG
While the information viewers use in describing a scene must evidently depend on what they see, previous studies have generally found that visual features are only weak predictors of what people tend to say. Beun and Cremers (
Attempts to elicit relational descriptions using scenes with target objects and a set of potential landmarks have had mixed results. Viethen and Dale (
Kelleher et al. (
In other words, while we know visual processing is tightly integrated with perception (Sedivy et al.,
This limited role for visual salience in production is the conclusion of a recent study (Gatt et al.,
Again, this REG result is puzzling in light of the extensive literature on perception (Eckstein,
We suspect that Gatt et al.'s result may reflect the types of images involved and will not necessarily generalize to more complex scenes. First, performing an exhaustive scan of a complex scene with hundreds or thousands of potential landmarks is prohibitively time-consuming. Secondly, the resulting descriptions, while guaranteed to be unambiguous, might refer to objects that listeners would nonetheless have great difficulty in finding. Lastly, the existence of completely separate visual mechanisms for perception and production (as opposed to for different types of scene) seems cognitively implausible. Although our study cannot rule out their proposal, it at least aims to establish an important role for visual salience in production as well as perception when the images involved are sufficiently complex.
Visual search and visual salience
Models of visual salience can be thought of as modeling two related mechanisms: low-level perceptual factors that render image regions more or less apparent, and the effect that these have on visual attention. Low-level models assign scores to pixels, or regions within a scene, that reflect their visual salience: how well they stand out from their surroundings. Over the past decade many different salience models have been developed by researchers in psychology, computer vision and robotics (see Toet,
A related field is that of visual search. In this paradigm, participants are presented with a stimulus and asked to decide, as quickly and accurately as they can, whether a pre-specified target is present or not. Stimuli typically consist of an array of shapes (although targets embedded within photographic stimuli are also used) and the challenge to researchers is to understand how the number of search items influences the difficulty of the task. The dominant theory is Guided Search (Wolfe,
When more naturalistic stimuli are used in visual search studies, there is no longer a simple way to represent the number of search items in the display. Instead, visual clutter (Rosenholtz et al.,
The Where's Wally images used in this study are certainly a favorable environment to find such effects. The Wally series is designed as a visual search game for children. The scenes are deliberately cluttered and contain large numbers of similar-looking people as well as more and less salient objects; in some sense they represent the other extreme to the simplistic scenes used in previous work. Results on such images certainly leave open a range of intermediate visual complexity in which salience effects might be weaker and harder to detect. But we would argue that the real world looks more complex than Figure 2. For example, over the 100 photographs used in Clarke et al.'s (submitted) object naming study, subjects were asked to look at each image, then look away and list all the objects they could recall. When the lists given by 24 subjects are reconciled, they contain a median of 26 objects per image. Spain and Perona (
Materials and methods
Data collection
A collection of 28 images taken from the Where's Wally picture books (Handford, 1987,
Participants (N = 155) were recruited via Amazon Mechanical Turk, a crowd-sourced marketplace (Munro et al.,
There was no time limit for either phase of the experiment. Participants took around 5 minutes on average to complete the task and were paid 40 cents. Data from three participants was excluded: two participants completed the task twice and a third participant returned a series of one-word referring expressions. The remaining 152 participants produced a dataset of 4256 descriptions. Of that larger dataset, the results reported here use 11 of the 28 scenes; this represents the subset of the data for which we have completed annotations and consists of 1672 descriptions (152 participants × 11 trials) over 176 targets (11 scenes × 16 targets).
Annotation
We annotated the elicited referring expressions to indicate which objects in the image were mentioned, which words in each expression referred to each object, and how the object references related to one another. Sample annotations are shown in examples (1) and (2). Words in <TARG> tags describe the target. Example (1) shows the annotated referring expression for an easy stimulus; a single landmark (the burning hut, indicated by the REL attribute) is used to localize the target. Example (2) shows the expression for a harder stimulus; two landmarks (the umbrella and ball) are introduced with the word “find” and marked with < EST> tags, and the ball is then used to localize the target. Objects in the image were labeled with bounding boxes (or for very large non-rectangular objects, bounding polygons). We did not distinguish references to geometrical parts of an object (“the left side of the track”) from references to the whole object, nor did we create separate boxes for small items that people wear or carry, or for architectural details of buildings (so “the boy in the yellow shirt” is treated as a single object). A few bounding boxes indicate groups of objects mentioned as a unit (“the three men”).
Example (1)
The < TARG> man < /TARG> just to the left of the < LMARK REL="TARG" OBJ="IMGID"> burning hut < /LMARK> < TARG> holding a torch and a sword < /TARG>.
Example (2)
Find < EST OBJ="IMGID1"> the red and white umbrella < /EST>. Then find < EST OBJ="IMGID2"> the blue and white beach ball < /EST>. Below and to the left < LMARK OBJ="IMGID2" REL="TARG"/> is < TARG> a dark skinned woman with a red bathing suit < /TARG>.
We marked the words in each expression which referred to or described each object and linked them to their corresponding bounding box. Words referring to the target were annotated with targ tags. When a reference to an object was used as a landmark in a relative description of another object (“the man just to the left of the burning hut”), we annotated it with an lmark ltag and indicated what object it was helping to locate.1 Objects mentioned without reference to another object (“find the X,” “there's an X”) were given an est tag (establish). When an expression picked out an individual without explicitly mentioning it, we created an empty phrase referring to it (so in example (2), “below and to the left” is annotated like “below and to the left of the ball”).
We validated our annotation scheme by independently annotating the elicited expressions for several targets in one image, then reconciling our results and updating the annotation guidelines. The authors of this paper contributed to the annotation of the referring expressions from 10 scenes, and the expressions from one additional scene were annotated by a paid annotator.
Visual features
For each scene, we assessed the salience of our annotated landmarks and targets, and the clutter of the scene as a whole. One complicating factor is that most salience models are based on the construction of a pixel-by-pixel salience map, and therefore they do not explicitly consider an object's area to be a contributing factor to how salient it is. Indeed, many salience models tend to undervalue large objects, as they contain large homogeneous regions. However, area is a basic visual property that should be considered in any common sense definition of what makes an object more or less salient, and therefore we will also include the square root of the area of a landmark's bounding box as a visual feature along with the pixel-based salience score (below). There is a significant correlation between them, r = 0.38.
To compute salience scores, we use the bottom-up component of Torralba et al.'s (
We measure the distance between each proposed landmark and the search target, computed between the closest points on their respective bounding boxes. Nearby objects were predicted to be better candidates for landmark mention. Finally, we also consider the visual clutter (feature congestion) of the scene, a measure that is related to the variability of features (color, orientation, and luminance) in a local neighborhood. Full details on measuring visual clutter can be found in Rosenholtz et al. (
Data transformations
The distributions of area and distance values in the dataset are skewed to the right. This is especially true for area; the dataset contains a few very large landmarks, while objects a corresponding number of deviations below the mean would have to have negative values. To counter this, we transform the values non-linearly. We use rather than area. This transformation is appropriate for several reasons (Gelman and Hill,
Analysis 1: length of expressions
There is a wide range in the length of the referring expressions in our dataset: between 1 and 104 words. As predicted, this variation appears to reflect the visual complexity of the scene (Figure 3): we find a correlation between the median length of referring expressions for targets in a scene and visual clutter (Spearman's rank correlation coefficient: p = 0.45, p = 0.02).3 We also see that the nature of referring expressions change as they get longer: short expressions typically only reference the target object, with an increasing number of landmarks being mentioned as the descriptions get longer (Figure 4).4
Figure 3

The median length of referring expressions for targets within an image varies with scene type (computed on all 28 images).
Figure 4

Longer referring expressions have proportionally fewer words describing the target (top plot shows median and quartiles) and proportionally more words describing other objects (bottom plot shows median and quartiles).
We model the length of expressions with both scene-level visual information (clutter) as well as object-level information (salience, area). We use these visual features to predict three outcomes: the total number of words in a description, the proportion of words referencing the target, and the number of landmarks mentioned. For the number of words and number of landmarks, we use linear mixed-effects models with a poisson linking function to model the count values. For the proportion of words referencing the target, we use a logistic mixed-effects model. All models contained factors for the salience and square root area of the target, the visual clutter of the scene, and interactions among those three factors.
Models were fit using the lmer function of the R package lme4 (Bates et al.,
For the overall number of words in the description, there was a main effect of area and marginal effects of salience and visual clutter (Table 1): descriptions were shorter for targets with larger area (β = −0.04) and greater salience (β = −0.03) and longer for targets in scenes with high clutter scores (β = 0.03). A marginal area × salience interaction indicated that these two negative effects are not quite additive, with a slightly reduced effect when both are large (β = 0.02). None of the other interactions reached significance.
Table 1
| β | SE | t-value | p-value | |
|---|---|---|---|---|
| Area | −0.04 | 0.02 | −2.29 | <0.05 |
| Salience | −0.03 | 0.02 | −1.76 | 0.08 |
| Clutter | 0.03 | 0.02 | 1.74 | 0.08 |
| Area × sal | 0.02 | 0.01 | 1.78 | 0.08 |
| Area × clutter | −0.02 | 0.02 | −0.92 | 0.36 |
| Sal × clutter | 0.01 | 0.02 | 0.76 | 0.45 |
| Area × sal × clutter | −0.01 | 0.02 | −0.27 | 0.79 |
Results of mixed-effects model for predicting number of overall words in a description.
Bolding indicates main effects or interactions that reached significance.
For the proportion of words referencing the target itself, there were main effects of area and salience (Table 2). Target descriptions were longer for those targets with larger area (β = 0.25) and greater salience (β = 0.20). There was no effect of clutter and the only interaction to reach significance was again the area × salience interaction, whereby the overall effect of these two factors is reduced when both are large (β = −0.11).
Table 2
| β | SE | z-value | p-value | |
|---|---|---|---|---|
| Area | 0.25 | 0.05 | 5.11 | <0.001 |
| Salience | 0.20 | 0.05 | 4.25 | <0.05 |
| Clutter | −0.02 | 0.04 | −0.52 | 0.60 |
| Area × sal | −0.11 | 0.04 | −2.78 | <0.01 |
| Area × clutter | 0.02 | 0.05 | 0.34 | 0.73 |
| Sal × clutter | 0.02 | 0.06 | 0.45 | 0.65 |
| Area × Sal × clutter | −0.04 | 0.05 | −0.57 | 0.57 |
Results of mixed-effects model for predicting proportion of words referencing the target in a description.
Bolding indicates main effects or interactions that reached significance.
For the number of landmarks included in the description, there were likewise effects of area and salience (Table 3). The number of landmarks mentioned decreased for targets with larger area (β = −0.14) and greater salience (β = −0.12). Again, area and salience interact (β = 0.07). Neither clutter nor any of the other interactions reached significance.
Table 3
| β | SE | z-value | p-value | |
|---|---|---|---|---|
| Area | −0.14 | 0.03 | −4.01 | <0.001 |
| Salience | −0.12 | 0.04 | −3.50 | <0.001 |
| Clutter | 0.00 | 0.03 | −0.13 | 0.89 |
| Area × Sal | 0.07 | 0.03 | 2.40 | <0.05 |
| Area × Clutter | −0.01 | 0.03 | −0.25 | 0.80 |
| Sal × Clutter | −0.05 | 0.04 | −1.25 | 0.21 |
| Area × Sal × Clutter | −0.02 | 0.04 | −0.56 | 0.58 |
Results of mixed-effects model for predicting number of landmarks included in a description.
Bolding indicates main effects or interactions that reached significance.
Analysis 2: choice of landmarks
The effects of an individual landmark's features on the probability of that landmark being chosen in an expression are shown in Figure 5. To measure the effect of visual properties on the choice to mention a particular landmark in a referring expression, we modeled the binary outcome of mention for each landmark in each description using a mixed-effects logistic regression. The model contained factors for the salience and square root area of the landmark, the distance between the landmark and the target, the visual clutter of the scene, and interactions among those four factors. Random participant-specific and target-specific intercepts and slopes were included (slopes were not crossed, due to the number of parameters to estimate and problems with model convergence). For this model, all objects in a scene were included, meaning that the mention outcome was 0 for most landmarks relative to most targets, since only a few landmarks were near enough or large/salient enough to merit mention. The set of ‘all objects’ consisted of every object that was mentioned in at least one referring expression in the dataset.
Figure 5

The effect of each feature on the conditional probability of naming a given object: (A) distance, (B) area, (C) salience. These figures were generated by dividing each feature range into 5% quantiles, and plotting the probability of mentioning a landmark from that quantile. Binomial distributions were then fitted to provide confidence intervals. The dotted line shows an estimate of the prior probability of selecting a landmark.
Again, we used the lmer function of the R package lme4. We report the coefficient estimates, standard error, and p-values based on the Wald Z statistic (Agresti,
The results (Table 4) show main effects of area, distance, and crucially visual salience: a landmark is more likely to be mentioned the larger it is (area: β = 0.57) and the more salient it is (salience: β = 0.25); it is less likely to be mentioned the farther it is from the target (distance: β = −0.99). The positive effect of area was stronger in more cluttered scenes (area × clutter: β = 0.20) and at greater distances (area × distance: β = 0.21), and this interaction with distance was stronger in more cluttered scenes (area × distance × clutter: β = 0.11). Again area and salience interact, reducing their overall effect when both are large (area × sal: β = −0.22), though this is less apparent in more cluttered scenes (area × salience × clutter interaction: β = −0.19) and at greater distances (area × salience × distance interaction: β = −0.09). Finally, the 4-way interaction was significant (area × salience × distance × clutter: β = −0.05), meaning that more distant, larger salient objects are less likely to be selected in cluttered scenes.
Table 4
| β | SE | p-value | |
|---|---|---|---|
| Lmark Area | 0.57 | 0.05 | <0.001 |
| Lmark Salience | 0.25 | 0.11 | <0.05 |
| Dist to targ | −0.99 | 0.05 | <0.001 |
| Clutter | 0.11 | 0.07 | 0.10 |
| Area × Sal | −0.22 | 0.04 | <0.001 |
| Area × Dist | 0.21 | 0.03 | <0.001 |
| Area × Clutter | 0.20 | 0.05 | <0.001 |
| Sal × Dist | 0.04 | 0.03 | 0.23 |
| Sal × Clutter | −0.03 | 0.11 | 0.78 |
| Dist × Clutter | 0.05 | 0.05 | 0.31 |
| Area × Sal × Dist | −0.09 | 0.02 | <0.001 |
| Area × Sal × Clut | −0.19 | 0.03 | <0.001 |
| Area × Dist × Clut | 0.11 | 0.03 | <0.001 |
| Sal × Dist × Clut | 0.00 | 0.03 | 0.99 |
| Area × Sal × Dist × Clut | −0.05 | 0.02 | <0.05 |
Results of mixed-effects model for predicting whether a landmark would be included in a description.
Bolding indicates main effects or interactions that reached significance.
Discussion
These results demonstrate that participants' production of referring expressions is affected by their perception of visual salience and clutter. As stated above, we agree with Viethen et al. (
Of course, salience is not the only driving force behind landmark selection. Participants might select landmarks with lower computed salience for a variety of reasons. In some cases, these landmarks appear to be intended as confirmation that the right object has been found, rather than an aid in finding the object to begin with. In others, their attention might be strongly directed toward the region around the target, so that objects appear perceptually salient to them despite not being visually salient to an observer who is unaware of the target's location. Such task-based effects on gaze and attentional allocation are known from other studies (Land et al.,
This raises the further question of how closely our computational salience prediction algorithm corresponds to actual human perception. Certainly it contributes something more than simple area and centrality (the model of salience implemented in Kelleher et al.,
While we have shown that salience has an effect on referring expression production, a critical question remains: do speakers choose to talk about salient objects in order to save themselves visual work, or do they perform a relatively comprehensive scan, but prefer to talk about objects that will be easier for listeners to find? In other words, is the observed effect driven by participant efficiency, or is it a case of “audience design” in which speakers try to make listeners' tasks efficient? REG models like the incremental algorithm of Pechmann (
The current study is insufficient to resolve this question. Although a negative finding (that visual salience had no effect) would have been fatal for the incremental model, the minimal-description model can incorporate visual salience (as in Kelleher et al.,
Beyond REG, our results also contribute to the ongoing debate surrounding the importance of salience in visual perception. Since the introduction of computational salience models, vision scientists have been able to test predictions from these models and compare them to the distributions of fixations obtained during eye-tracking studies. Specifically, the majority of this work has centered around the question of whether bottom-up salience can provide a robust explanation for the distribution of fixation locations during a variety of tasks such as free-viewing, visual search, and scene memorization. Furthermore, bottom-up salience is frequently taken as a benchmark to evaluate other factors against. For example, Tatler (
The work presented here shows that low-level visual salience plays an important role even in higher-level task-driven cognitive behavior. However, results like these suggest that a more object-centric model of visual attention might do even better. Our results support the idea of a close connection between vision and language, where relatively low-level mechanisms on one side can influence the other. We hope that further study of tasks like REG can reveal more about this interface and what kinds of information pass through it.
This study shows a clear effect of visual properties on the production of referring expressions, both in length and in composition. This conclusion may seem obvious – surely the complexity of the image people are looking at should affect what they say. But nonetheless, over a decade of research has failed to meaningfully establish it, producing instead a confusing array of weak results and failures to find significance. Moreover, this gap in the research record has had significant influence on the models proposed for REG. Psychological models like Gatt et al. (
This paper should serve as to correct such views. For sufficiently complex images, visual features do matter, and the coefficients in our models make explicit predictions about how much a particular degree of corpus-wide variation in visual salience is expected to influence the results. In order to generalize to the full range of human performance, we argue that future models of REG should incorporate up-to-date models of low-level perception from the vision literature. Their performance should be evaluated on complex images with hundreds of objects, each differing in salience, as well as the arrays of ten or twenty similar-looking objects used in previous work. Finally, vision scientists working on salience should consider their models to be more than simple fixation predictors; visual salience has high-level cognitive effects which surface even in simple experiments.
Conflict of interest statement
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Statements
Acknowledgments
This research was supported by EPSRC grant EP/H050442/1 and European Research Council grant 203427 “Synchronous Linguistic and Visual Processing.” We also thank Ellen Bard, Frank Keller, Robin Hill, and Meg Mitchell for helpful discussion and support and Louisa Miller for help in annotation. Jette Viethen for the use of one of her figures, and our reviewers, Piers Howe and Joseph Schmidt.
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Footnotes
1.^If it was unclear which object was the main object of the description and which the landmark, we preferred to make the target the main object. Then any landmarks that were directly relative to the target, and so on. This is because the participant's task was to refer to the target; other objects were presumably introduced to help locate it and not vice versa.
2.^We also investigated log (l + area), but the linear correlation with landmark choice is not as strong.
3.^Because this analysis relies only on the total number of words, regardless of what they describe, we perform the analysis on all 28 scenes. On the 11 annotated images, we obtain a similar result (p = 0.47) but with less power (p = 0.14).
4.^Although the number of words devoted to targets decreases as a proportion of the total as expressions lengthen, the absolute number of words devoted to targets increases.
References
1
AgrestiA. (2002). Categorical Data Analysis, 2nd Edn. Hoboken: Wiley.
2
AsherM. F.TolhurstD. J.TrosciankoT.GilchristI. D. (2013). Regional effects of clutter on human target detection performance. J. Vis. 10.1167/13.5.25
3
BaayenR.DavidsonD.BatesD. (2008). Mixed-effects modeling with crossed random effects for subjects and items. J. Mem. Lang. 59, 390–412. 10.1016/j.jml.2007.12.005
4
BardE.HillR.AraiM. (2009). Referring and gaze alignment: accessibility is alive and well in situated dialogue, in Proceedings of the 31st Annual Meeting of the Cognitive Science Society, Amsterdam, 1246–1251.
5
BatesD.MaechlerM.BolkerB. (2011). lme4: linear mixed-effects models using S4 classes. R package verson 0.999375-32.
6
BeunR.-J.CremersA. H. (1998). Object reference in a shared domain of conversation. Pragmat. Cogn. 6, 121–152. 10.1075/pc.6.1-2.08beu
7
Brown-SchmidtS.TanenhausM. (2008). Real-time investigation of referential domains in unscripted conversation: a targeted language game approach. Cogn. Sci. 32, 643–684. 10.1080/03640210802066816
8
ClarkH.Wilkes-GibbsD. (1986). Referring as a collaborative process. Cognition22, 1–39. 10.1016/0010-0277(86)90010-7
9
CocoM. I.KellerF. (2012). Scan pattern predicts sentence production in the cross-modal processing of visual scenes. Cogn. Sci. 36, 1204–1223. 10.1111/j.1551-6709.2012.01246.x
10
DaleR.ReiterE. (1995). Computational interpretations of the Gricean maxims in the generation of referring expressions. Cogn. Sci. 19, 233–263. 10.1207/s15516709cog1902_3
11
EcksteinM. P. (2011). Visual search: a retrospective. J. Vis. 11, 1–36. 10.1167/11.5.14
12
EinhauserW.SpainM.PeronaP. (2008). Objects predict fixations better than early saliency. J. Vis. 8, 1–26. 10.1167/8.14.18
13
Fei-FeiL.IyerA.KochC.PeronaP. (2007). What do we perceive in a glance of a real-world scene?J. Vis. 7, 1–29. 10.1167/7.1.10
14
GattA.van GompelR. P. G.KrahmerE.van DeemterK. (2012). Does domain size impact speech onset time during reference production? in Proceedings of the 34th Annual Meeting of the Cognitive Science Society, Sapporo, 1584–1589.
15
GelmanA.HillJ. (2007). Data Analysis Using Regression and Multilevel/Hierarchical Models, Volume Analytical Methods for Social Research. New York: Cambridge University Press.
16
HandfordM. (1987). Where's Wally?, 3rd Edn. London: Walker Books.
17
HandfordM. (1988). Where's Wally Now?4th Edn. London: Walker Books.
18
HandfordM. (1993). Where's Wally?, 3rd Edn. London: Walker Books.
19
HendersonJ.ChanceauxM.SmithT. (2009). The influence of clutter on real-world scene search: evidence from search efficiency and eye movements. J. Vis. 9, 1–8. 10.1167/9.1.32
20
HortonW.KeysarB. (1996). When do speakers take into account common ground?Cognition59, 91–117. 10.1016/0010-0277(96)81418-1
21
IttiL.KochC. (2000). A saliency-based search mechanism for overt and covert shifts of visual attention. Vision Res. 40, 1489–1506. 10.1016/S0042-6989(99)00163-7
22
KelleherJ.CostelloF.van GenabithJ. (2005). Dynamically structuring, updating and interrelating representations of visual and linguistic discourse context. Artif. Intell. 167, 62–102. 10.1016/j.artint.2005.04.008
23
KrahmerE.van DeemterK. (2012). Computational generation of referring expressions: a survey. Comput. Ling. 38, 173–218. 10.1162/COLI_a_00088
24
LandM.MennieN.RustedJ. (1999). The roles of vision and eye movements in the control of activities of daily living. Perception28, 1311–1328. 10.1068/p2935
25
LouwerseM.BeneshN.HoqueM.JeuniauxP.LewisG.WuJ.et al. (2007). Multimodal communication in face-to-face computer-mediated conversations, in Proceedings of the 29th Annual Meeting of the Cognitive Science Society, Mahwah, 1235–1240.
26
MitchellM.van DeemterK.ReiterE. (in press). Generating expressions that refer to visible objects, in Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Atlanta: Association for Computational Linguistics).
27
MunroR.BethardS.KupermanV.LaiV. T.MelnickR.PottsC.et al. (2010). Crowdsourcing and language studies: the new generation of linguistic data, in Proceedings of the NAACL HLT 2010 Workshop on Creating Speech and Language Data with Amazon's Mechanical Turk (Los Angeles: Association for Computational Linguistics), 122–130.
28
NuthmannA.HendersonJ. M. (2010). Object-based attentional selection in scene viewing. J. Vis. 10, 1–19. 10.1167/10.8.20
29
PechmannT. (2009). Incremental speech production and referential overspecification. Linguistics27, 89–100. 10.1515/ling.1989.27.1.89
30
RosenholtzR.LiY.NakanoL. (2007). Measuring visual clutter. J. Vis. 7, 1–21. 10.1167/7.2.17
31
SedivyJ. C. (2003). Pragmatic versus form-based accounts of referential contrast: evidence for effects of informativity expectations. J. Psycholinguist. Res. 32, 3–23. 10.1023/A:1021928914454
32
SedivyJ. C.TanenhausM. K.ChambersC. G.CarlsonG. N. (1999). Achieving incremental semantic interpretation through contextual representation. Cognition71, 109–147. 10.1016/S0010-0277(99)00025-6
33
SimoncelliE.FreemanW. (1995). The steerable pyramid: a flexible architecture for multi-scale derivative computation, in Proceedings, International Conference on Image Processing 1995, Vol. 3 (Washington: IEEE), 444–447.
34
SpainM.PeronaP. (2010). Measuring and predicting object importance. Int. J. Comput. Vis. 91, 59–76. 10.1007/s11263-010-0376-0
35
TanenhausM. K.SpiveyM. J.EberhardK. M.SedivyJ. C. (1995). Integration of visual and linguistic information in spoken language comprehension. Science268, 1632–1634. 10.1126/science.7777863
36
TatlerB. W. (2007). The central fixation bias in scene viewing: selecting an optimal viewing position independently of motor biases and image feature distributions. J. Vis. 7, 1–17. 10.1167/7.14.4
37
TatlerB. W.MelcherD. (2007). Pictures in mind: initial encoding of object properties varies with the realism of the scene stimulus. Perception36, 1715–1729. 10.1068/p5592
38
ToetA. (2011). Computational versus psychophysical bottom-up image saliency: A comparative evaluation study. IEEE Trans. Pattern Anal. Mach. Intell. 33, 2131–2146. 10.1109/TPAMI.2011.53
39
TorralbaA.OlivaA.CastelhanoM.HendersonJ. M. (2006). Contextual guidance of attention in natural scenes: The role of global features on object search. Psychol. Rev. 113, 766–786. 10.1037/0033-295X.113.4.766
40
ViethenJ.DaleR.GuheM. (2011). The impact of visual context on the content of referring expressions, in Proceedings of the 13th European Workshop on Natural Language Generation (Nancy: Association for Computational Linguistics), 44–52.
41
ViethenJ.DaleR. (2006). Algorithms for generating referring expressions: do they do what people do? in Proceedings of the Fourth International Natural Language Generation Conference, INLG '06 (Stroudsburg, PA: Association for Computational Linguistics), 63–70.
42
ViethenJ.DaleR. (2008). The use of spatial relations in referring expressions, in Proceedings of the 5th International Conference on Natural Language Generation.
43
ViethenJ.DaleR. (2011). GRE3D7: “A corpus of distinguishing descriptions for objects in visual scenes,” in Proceedings of the Workshop on Using Corpora in Natural Language Generation and Evaluation, Edinburgh.
44
WolfeJ. M. (1994). Guided search 2.0: a revised model of visual search. Psychon. Bull. Rev. 1, 202–238. 10.3758/BF03200774
45
WolfeJ. M. (2012). Visual search, in Cognitive Search: Evolution, Algorithms and the Brain, eds ToddP.HollsT.RobbinsT. (Cambridge: MIT Press), 159–175.
Summary
Keywords
referring expression generation, visual salience, visual clutter
Citation
Clarke ADF, Elsner M and Rohde H (2013) Where's Wally: the influence of visual salience on referring expression generation. Front. Psychol. 4:329. doi: 10.3389/fpsyg.2013.00329
Received
11 March 2013
Accepted
21 May 2013
Published
18 June 2013
Volume
4 - 2013
Edited by
Tamara Berg, Stony Brook University, USA
Reviewed by
Piers D. L. Howe, Harvard Medical School, USA; Joseph Schmidt, University of South Carolina, USA
Copyright
© 2013 Clarke, Elsner and Rohde.
This is an open-access article distributed under the terms of the Creative Commons Attribution License, which permits use, distribution and reproduction in other forums, provided the original authors and source are credited and subject to any copyright notices concerning any third-party graphics etc.
*Correspondence: Micha Elsner, Department of Linguistics, The Ohio State University, 1712 Neil Avenue, Columbus, OH 43210, USA e-mail: melsner@ling.osu.edu
This article was submitted to Frontiers in Perception Science, a specialty of Frontiers in Psychology.
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.