Abstract
Three times per second, our eyes make a new fixation that generates a new bottom-up analysis in the visual system. How much is extracted from each glimpse? For how long and in what form is that information remembered? To answer these questions, investigators have mimicked the effect of continual shifts of fixation by using rapid serial visual presentation of sequences of unrelated pictures. Experiments in which viewers detect specified target pictures show that detection on the basis of meaning is possible at presentation durations as brief as 13 ms, suggesting that understanding may be based on feedforward processing, without feedback. In contrast, memory for what was just seen is poor unless the viewer has about 500 ms to think about the scene: the scene does not need to remain in view. Initial memory loss after brief presentations occurs over several seconds, suggesting that at least some of the information from the previous few fixations persists long enough to support a coherent representation of the current environment. In contrast to marked memory loss shortly after brief presentations, memory for pictures viewed for 1 s or more is excellent. Although some specific visual information persists, the form and content of the perceptual and memory representations of pictures over time indicate that conceptual information is extracted early and determines most of what remains in longer-term memory.
Introduction
The problem
We make three or four eye fixations each second, all day long. That suggests that 250 ms is long enough to identify most objects, but is it enough to recognize a whole scene? How much do we remember about each fixation and for how long? To develop and maintain information about the environment, we need some form of visual memory that spans several fixations. But, carryover from the preceding fixation lacks detail (e.g., Irwin, ; Irwin and Andrews, ; Henderson and Hollingworth, ). Indeed, we overlook major changes in a scene if the scene is interrupted for as little as 80 ms – the phenomena of change blindness (e.g., Rensink et al., , ) and boundary extension (Intraub and Richardson, ). We are not blind, however, to changes that affect gist or changes to objects that we are attending or are about to fixate. Thus, the information that we carry over from a fixation seems to be limited and to be meaningful rather than purely visual. Our memory for pictures is poor, however, for unrelated pictures presented in a continuous sequence at rates in the range of eye fixations (Potter and Levy, ).
On the other hand, we have good long-term memory for pictures viewed for 1 or 2 s (Nickerson, ; Shepard, ; Potter and Levy, ; Standing, ) and we can remember them in considerable detail, whether they represent single objects (Brady et al., ) or more complex scenes (Konkle et al., ). Some highly distinctive information must be retained from pictures that we have viewed for a few seconds.
How long does it take to understand a visual object or scene?
There is no simple answer to this question: the answer depends on what one’s criterion for understanding is. One could measure the time it takes to name the picture, but even for objects with well-known names, the naming time of about 900 ms includes search for the word after one has already recognized what the object is. Another measure of understanding is the time to decide whether the scene or object matches some description, such as “animal.” This category detection task turns out to be considerably faster (around 600 ms) than the time to name a picture (Potter and Faulconer, ), but still includes the time to generate the yes or no response. A still faster response is the time between the onset of a pair of pictures and the initiation of an eye movement to (for example) the picture of an animal or a face (e.g., Kirchner and Thorpe, ; Crouzet et al., ): this selective decision can take as little as 100 ms. All the measures just discussed include the time for the information to pass from the retina to the visual cortex as well as decision and response processes that occur after identification (e.g., Potter, ). Still shorter times can be obtained by using measures of brain responses such as event related potentials (ERPs) that do not include any overt response.
Single masked stimuli
A different approach is to control the time available for processing a stimulus such as a picture, and to measure the minimum presentation time required for successful identification. However, because of visual persistence (continued activation in the visual system after a stimulus ends), the duration of the physical stimulus is not closely related to the effective duration of the stimulus: for example, a picture presented for only 20 ms followed by a blank screen may be as readily processed as one shown for 100 ms. A common method to solve that problem is to use a backward mask such as a patterned stimulus that follows the picture. Such a mask is thought to interrupt processing of the picture. If that is the case, then by varying the stimulus onset asynchrony (SOA) between the onset of the target and that of the mask, a minimal processing time required for identification can be determined. For example, when a single picture is presented, followed by a visual mask (such as a collage of colored paper cut into small circles and irregular shapes), it is possible to remember as many as half the pictures with a duration as short as 50 ms, and 80% are remembered at a duration of 120 ms (Potter, ; see Figure 3, discussed below).
A continuing problem with the logic of the masking procedure, however, is that the neural basis for the effect is not well-understood: does the masked stimulus continue to be processed, perhaps unconsciously, after the mask appears, or does processing instantly stop? This question is especially relevant to feedforward models, discussed below. Moreover, a backward mask does not necessarily interrupt all processing – and the amount of interference is a complex function of the visual relation between the target and mask, the semantic (conceptual) relation, and the SOA between target and mask. (For a more complete account of the complexities of backward and forward masking, see Eriksen and Eriksen, , and Breitmeyer and Ogmen, .) With very short SOAs, the visual relation may be a stronger determinant of the effectiveness of the mask than the conceptual relation, but as the SOA increases, the reverse may be the case (e.g., Potter, ; Loftus and Ginn, ). Indeed, the most important factor may be whether the following mask is itself a stimulus that the viewer needs to attend to and report on: see rapid serial visual presentation (RSVP) below and the discussion of visual versus conceptual masking. I return to the question of masking in Section “Detecting Pictures at Ultra-High Rates: Evidence for Feedforward Processing?”
Perception of objects in settings
A further question is whether knowledge of co-occurrences between objects and settings influences the initial perception of a scene, or whether (as suggested by Hollingworth and Henderson, , ) objects and settings in a given picture are first understood independently and only later merged. In one set of studies (Davenport and Potter, ), pictured objects such as a football player or a priest were superimposed, either congruently or incongruently, on background settings such as a football field or the interior of a cathedral (Figure 1). The pictures were presented for 80 ms, with a backward noise mask, and the participant was instructed to report the foreground object, the background setting, or both. In each case performance was better in the congruent than the incongruent condition, suggesting that objects and background are processed interactively, early in processing. In a further study (Davenport, ) one or two objects were presented on a background. The relation between the two objects (whether they would be likely to be present in the same scene or not) had an effect on report that was additive with the effect of congruency with the background: that is, the relationship between the two objects, as well as each of the objects’ relation to the background, influenced report of the objects. Joubert et al. (, ) carried out similar studies, finding that objects in congruent contexts were responded to faster than in incongruous contexts.
Figure 1
Rapid Serial Visual Presentation
In studies using backward masking of pictures, each trial consists of a single picture and a mask. In normal vision, however, the eyes make a continuous sequence of fixations. What happens when pictures are presented in a continuous stream at durations in the range of eye fixations, and participants try to remember all of them? To investigate this question, Potter and Levy (
Figure 2

An illustration of an RSVP sequence of pictures.
Figure 3

Proportion of pictures recognized following single masked presentations (solid curve, Potter,
Visual versus conceptual masking
What makes an RSVP sequence hard to remember is not the briefness of the pictures, but the fact that each picture is immediately followed by another. With a single masked picture, the viewer can continue to process the information after the mask appears; evidently that is not possible with a continuous sequence in which all the pictures need to be attended. In a study by Intraub (
Once the SOA between the picture and the following visual mask is about 100 ms, memory depends little on the actual duration of presentation, but instead on the total uninterrupted time the viewer has to continue to think about the picture. Thus, if a viewer is shown a sequence of pictures that alternate between a short duration of 112 ms and a long duration of 1500 ms, the instruction to attend only to the brief pictures results in memory for about 63% of the brief pictures and only 54% of the long pictures: intention to continue processing the brief pictures actually leads to better memory than for the long-duration pictures (Intraub,
Detecting pictures
Given the poor memory for pictures presented at durations in the range of eye fixations, does it take longer than a single fixation to understand a novel scene? Do we even momentarily understand pictures shown for only 250 ms in an RSVP sequence? Intuitively, we may think that if we had understood what a picture was about, we would surely remember it for at least a few minutes. Perhaps viewers fail to remember briefly presented pictures because they did not comprehend them in a single glimpse, whereas normally they could continue to look at something until it is recognized. Yet, when viewing pictures at a rate as high as 10/s, one’s impression is that each picture can be seen and understood momentarily: is that an illusion? To address this question, participants were asked to detect a target picture in an RSVP sequence that was named or shown to them before the sequence (Potter,
Figure 4

Detection of a target picture in an RSVP sequence of 16 pictures, given a picture of the target or a name for the target, as a function of the presentation time per picture. Also shown is later recognition performance in a group that simply viewed the sequence, and then was tested for recognition. Results are corrected for guessing (see text footnote 1). From Potter (
Further evidence that a picture’s identity can be retrieved quickly is shown in a detection study (Potter et al.,
Figure 5

An example of an RSVP sequence in a search experiment in which participants reported the specific names of two exemplars of the search category. Here the exemplars are hamburger and spaghetti. From Potter et al. (
Detection and memory when multiple pictures are presented simultaneously
Viewers can process serially presented pictures remarkably rapidly, but can they process two or more pictures presented simultaneously? When the task is to detect a specified target, the results suggest that detection is relatively successful with up to four simultaneous pictures, in RSVP streams consisting of eight successive four-item arrays (Potter and Fox,
Rapid memory loss for pictures seen briefly in RSVP: Serial position effects in memory testing
People can understand pictures presented briefly, but forget most of them a few minutes later. When the recognition test begins immediately, the first one or two pictures tested are likely to be remembered well, but there is rapid loss over the next several seconds of testing (Potter et al.,
What is the nature of this short-lasting memory for pictures?
The time course of forgetting after viewing an RSVP sequence of pictures contrasts with that of change blindness, the apparently immediate loss of detailed information about a single picture, once it is no longer in view. Change blindness is the inability of viewers to detect a change in one feature of a picture, and it has been observed when a blank interval as short as 80 ms intervenes between the initial and changed versions; at longer intervals, the problem is even more acute (see Rensink et al.,
Could the short-lasting memory for pictures be iconic memory (e.g., Sperling,
A likely contributor to short-term memory for pictures is conceptual short-term memory (CSTM), a short-lasting memory component proposed by Potter (
Conceptual versus visual-perceptual memory
In relation to rapidly presented pictures, the CSTM claim is that some pictures are adequately encoded and consolidated into longer-term memory during even brief viewing, but others are represented only in CSTM and are vulnerable to interference in the first few seconds after viewing. However, we do not know whether the picture representation that persists for several seconds in the studies we have reviewed here is sufficiently abstract to be considered conceptual rather than wholly or partly perceptual. Do viewers remember only the picture’s conceptual content or gist, or do they also remember visual features such as color, shape, and layout? Work of Irwin and Andrews (
The relative roles of such specific pictorial information and more abstract conceptual information were explored in Potter et al. (
Figure 6

Recognition test of five pictures shown in RSVP for 173 ms/picture; the test used pictures or titles. Guessing-corrected results (see text footnote 1) are shown as a function of relative position in the recognition test, which included five new pictures (distractors). From Potter et al. (
In a further test of the conceptual basis of memory, Potter et al. (
Short-lasting memory: Summary
In sum, initial memory for a glimpsed picture (seen for the equivalent of a single fixation) is fairly accurate, but declines markedly over the first few recognition tests (or across an unfilled delay of 5 s). There is some evidence that the initial stronger memory includes specifically visual information, whereas after a delay the memory is primarily conceptual. That is, detailed visual information about a picture is lost more rapidly than conceptual information. Accurate visual information may be important for maintaining and updating scene representations over fixations, but conceptual memory seems to be the basis for longer-term, organized knowledge.
As stated earlier, unlike briefly glimpsed pictures, memory for pictures viewed for a second or more can be highly accurate, at least when viewers are paying attention. Yet, as work reviewed here shows, normal eye fixations are too brief to guarantee good memory. They are, however, long enough to make it highly likely that the viewer will have understood what he or she saw, at least momentarily, allowing the viewer to continue looking or to take appropriate action. The rapid comprehension of the gist of a scene suggests that scenes are initially perceived as wholes – like single objects. Although the gist of pictured scenes can be extracted rapidly, exactly how that is done remains unclear. Work of Oliva and her collaborators has given us some ideas about how visual properties such as layout, texture, color, and the like can enable rapid categorization of natural scenes, street scenes, and interiors (Oliva,
Detecting Pictures at Ultra-High Rates: Evidence for Feedforward Processing?
What constitutes evidence for feedforward processing?
It is widely assumed that under normal viewing conditions perception results from a combination of feedforward and feedback connections (Di Lollo et al.,
Conscious perception
The ability to identify or remember a stimulus is commonly taken to mean that the viewer was conscious of the stimulus, and in the work discussed here I make the assumption that consciousness is shown by the ability to report on the stimulus by responding to a target picture or by recognizing its title or the picture itself, in a memory test. (See, however, evidence for unconscious effects, in Feedforward Processing and Masked Priming.) There is a debate about whether a single forward pass is sufficient for conscious perception. A reentrant process providing feedback may be necessary to achieve understanding and conscious awareness (Lamme and Roelfsema,
Evidence for processing of very brief stimuli
RSVP responses: monkey neurons and humans
Recordings of individual neurons in the cortex of the anterior superior temporal sulcus (STSa) of monkeys who viewed a set of pictures of monkey faces and other objects via RSVP at various rates up to 72/s (14 ms) showed that neurons respond to a preferred picture above chance, even at 14 ms (Keysers et al.,
Further evidence: detection and immediate memory
A study by Potter, Wyble, and McCourt (in preparation) replicated some of the behavioral conditions of Keysers et al. (
These results are consistent with the claim of the feedforward model that pictures can be understood in a single feedforward sweep even when attention has not been directed to a specific category in advance.
How long does recognition memory last, after a very brief presentation?
Studies of the monkey visual system using single-cell recordings show that cortical neurons that are selective for particular objects can “recognize” multiple objects in parallel at levels as high as the inferior temporal cortex. Something similar in human perception might account for the ability to remember rapidly presented pictures. In monkeys, this initial parallel process is followed within 150 ms by competitive inhibition of all but the one relevant object in a given receptive field, at least when there is a task that defines the relevant stimulus (e.g., Chelazzi et al.,
Feedforward processing and masked priming
In masked priming studies, a brief presentation of a word becomes invisible when it is followed by a second unmasked word to which the participant must respond (Forster and Davis,
Discussion: Ultra-rapid processing and feedforward processing
Both the results of Keysers et al. (
But are there other explanations for successful detection when the presentation duration is brief and masked by successive pictures? One possibility is that at high rates of presentation several temporally adjacent pictures are integrated, like double or multiple camera exposures. Certainly the subjective impression in viewing rapid sequences is that the pictures merge into each other visually, as though they were overlaid. Possibly viewers simply recover the target from such a composite representation, rather than detecting it during the feedforward pass.
A related possibility is that following masks do not interrupt processing immediately. As mentioned in Section “Single Masked Stimuli,” the neural basis for masking is not well-understood. There is evidence (Keysers et al.,
Nonetheless, the feedforward hypothesis remains a strong contender as an explanation of picture identification with very brief presentation durations. In the absence of a specific model for how feedback might assist reportable detection of brief targets, the feedforward hypothesis seems the most plausible account.
How Long Does it Take to Understand a Pictured Scene?
Returning to the question considered in the introduction, what can be concluded about the time required to identify a scene? If the question is the minimum exposure duration (prior to a mask) that is required, 13 or 14 ms is sometimes enough, when the mask is another scene. But if the question is the time from arrival at the retina to correct categorization, then the most reliable measures available at present are reaction time measures, the most sensitive of which is an eye movement to the appropriate target in a choice situation. For detection of a face (when a picture with a face is presented together with another picture), that time can be as short as 100 ms, with a mean time of 140 ms (Crouzet et al.,
Detection of a pre-specified target is likely to be more rapid than comprehension of a new scene that the viewer is told nothing about. That is one reason that recognition memory for a picture, even immediately after a short RSVP sequence, is less accurate than detection (Keysers et al.,
As reviewed above, the speed of detection has suggested to a number of investigators that accurate comprehension or categorization can occur on the basis of the early feedforward sweep of visual information, without requiring feedback loops from higher to lower levels and back. The behavior of individual neurons in the human inferior temporal cortex and homologous areas in monkey cortex, reviewed elsewhere in this issue, provides another window on the timing of picture categorization that gives some support to the feedforward hypothesis. When a scene is complex or its components are unfamiliar, we undoubtedly require more processing time and probably more than a single fixation to comprehend it. However, a lifetime of knowledge of the world that is built into our visual system appears to allow immediate understanding of most scenes, based on the initial sweep of visual information when the scene is presented.
Statements
Acknowledgments
This work was supported by Grant MH47432 from the National Institute of Mental Health.
Conflict of interest
The author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Footnotes
1.^A one-high-threshold formula was used to correct for guessing, Pcorr = [P(TY) − P(FY)]/[1 − P(FY)], where TY is a correct yes response and FY is a false yes response. This guessing correction is used in all data figures.
References
1
BradyT. F.KonkleT.AlvarezG. A.OlivaA. (2008). Remembering thousands of objects with high fidelity. Proc. Natl. Acad. Sci. U.S.A.105, 14325–14329.10.1073/pnas.0803390105
2
BreitmeyerB. G.OgmenH. (2006). Visual Masking: Time Slices through Conscious and Unconscious Vision, 2nd Edn. New York: Oxford University Press.
3
ChelazziL.DuncanJ.MillerE. K.DesimoneR. (1998). Responses of neurons in inferior temporal cortex during memory-guided visual search. J. Neurophysiol.80, 2918–2940.
4
ColtheartM. (1983). Iconic memory. Philos. Trans. R. Soc. Lond. B Biol. Sci.302, 283–294.10.1098/rstb.1983.0055
5
CrouzetS. M.KirchnerH.ThorpeS. J. (2010). Fast saccades toward faces: face detection in just 100 ms. J. Vis.10, 16.1–17.10.1167/10.4.16
6
DavenportJ. L. (2007). Consistency effects between objects in scene processing. Mem. Cognit.35, 393–401.10.3758/BF03193280
7
DavenportJ. L.PotterM. C. (2004). Scene consistency in object and background perception. Psychol. Sci.15, 559–564.10.1111/j.0956-7976.2004.00719.x
8
DehaeneS.NaccacheL. (2001). Towards a cognitive neuroscience of consciousness: basic evidence and a workspace framework. Cognition79, 1–37.10.1016/S0010-0277(00)00123-2
9
DeheaneS.KergsbergM.ChangeuxJ. P. (1998). A neuronal model of a global workspace in effortful cognitive tasks. Proc. Natl. Acad. Sci. U.S.A.95, 14529–14534.10.1073/pnas.95.24.14529
10
Del CulA.BailletS.DehaeneS. (2007). Brain dynamics underlying the nonlinear threshold for access to consciousness. PLoS Biol.5, 2408–2423.10.1371/journal.pbio.0050260
11
Di LolloV.EnnsJ. T.RensinkR. A. (2000). Competition for consciousness among visual events: the psychophysics of reentrant visual pathways. J. Exp. Psychol. Gen.129, 481–507.10.1037/0096-3445.129.4.481
12
EndressA. D.PotterM. C. (in press). Early conceptual and linguistic processes operate in independent channels. Psychol. Sci. [Epub ahead of print].
13
EnnsJ. T.Di LolloV. (2000). What’s new in visual masking?Trends Cogn. Sci. (Regul. Ed.)4, 345–352.10.1016/S1364-6613(00)01520-5
14
EriksenC. W.EriksenB. A. (1971). Visual perceptual processing rates and backward and forward masking. J. Exp. Psychol.89, 306–313.10.1037/h0031160
15
ForsterK. I. (1970). Visual perception of rapidly presented word sequences of varying complexity. Percept. Psychophys.8, 215–221.10.3758/BF03210208
16
ForsterK. I.DavisC. (1984). Repetition priming and frequency attenuation in lexical access. J. Exp. Psychol. Learn. Mem. Cogn.10, 680–698.10.1037/0278-7393.10.4.680
17
GordonR. D.IrwinD. E. (2000). The role of physical and conceptual properties in preserving object continuity. J. Exp. Psychol. Learn. Mem. Cogn.26, 136–150.10.1037/0278-7393.26.1.136
18
HendersonJ. M. (1997). Transsaccadic memory and integration during real-world object perception. Psychol. Sci.8, 51–55.10.1111/j.1467-9280.1997.tb00543.x
19
HendersonJ. M.HollingworthA. (1999). High-level scene perception. Annu. Rev. Psychol.50, 243–271.10.1146/annurev.psych.50.1.243
20
HochsteinS.AhissarM. (2002). View from the top: hierarchies and reverse hierarchies in the visual system. Neuron36, 791–804.10.1016/S0896-6273(02)01091-7
21
HollingworthA.HendersonJ. M. (1998). Does consistent scene context facilitate object perception?J. Exp. Psychol. Gen.127, 398–415.10.1037/0096-3445.127.4.398
22
HollingworthA.HendersonJ. M. (1999). Object identification is isolated from scene semantic constraint: evidence from object type and token discrimination. Acta Psychol. (Amst.)102, 319–343.10.1016/S0001-6918(98)00053-5
23
HungC. P.KreimanG.PoggioT.DiCarloJ. J. (2005). Fast readout of object identity from macaque inferior temporal cortex. Science310, 863–866.10.1126/science.1116739
24
IntraubH. (1979). The role of implicit naming in pictorial encoding. J. Exp. Psychol. Hum. Learn.5, 1–12.10.1037/0278-7393.5.2.78
25
IntraubH. (1980). Presentation rate and the representation of briefly glimpsed pictures in memory. J. Exp. Psychol. Hum. Learn.6, 1–12.10.1037/0278-7393.6.1.1
26
IntraubH. (1981). Rapid conceptual identification of sequentially presented pictures. J. Exp. Psychol. Hum. Percept. Perform.7, 604–610.10.1037/0096-1523.7.3.604
27
IntraubH. (1984). Conceptual masking: the effects of subsequent visual events on memory for pictures. J. Exp. Psychol. Learn. Mem. Cogn.10, 115–125.10.1037/0278-7393.10.1.115
28
IntraubH.RichardsonM. (1989). Wide-angle memories of close-up scenes. J. Exp. Psychol. Learn. Mem. Cogn.15, 1989, 179–187.10.1037/0278-7393.15.2.179
29
IrwinD. E. (1992). Memory for position and identity across eye movements. J. Exp. Psychol. Learn. Mem. Cogn.18, 307–317.10.1037/0278-7393.18.2.307
30
IrwinD. E.AndrewsR. V. (1996). “Integration and accumulation of information across saccadic eye movements,” in Attention and Performance XVI: Information Integration in Perception and Communication, eds InuiT.McClellandJ. L. (Cambridge, MA: MIT Press), 125–155.
31
JoubertO.FizeD.RousseletG. A.Fabre-ThorpeM. (2008). Early interference of context congruence on object processing in rapid visual categorization of natural scenes. J. Vis.8, 11.1–11.18.10.1167/8.13.11
32
JoubertO. R.RousseletG. A.FizeD.Fabre-ThorpeM. (2007). Processing scene context: fast categorization and object interference. Vision Res.47, 3286–3297.10.1016/j.visres.2007.09.013
33
KeysersC.XiaoD. K.FöldiákP.PerrettD. I. (2001). The speed of sight. J. Cogn. Neurosci.13, 90–101.10.1162/089892901564199
34
KeysersC.XiaoD.-K.FöldiákP.PerrettD. I. (2005). Out of sight but not out of mind: the neurophysiology of iconic memory in the superior temporal sulcus. Cogn. Neuropsychol.22, 316–332.10.1080/02643290442000103
35
KirchnerH.ThorpeS. J. (2006). Ultra-rapid object detection with saccadic eye movements: visual processing speed revisited. Vision Res.46, 1762–1776.10.1016/j.visres.2005.10.002
36
KonkleT.BradyT. F.AlvarezG. A.OlivaA. (2010). Scene memory is more detailed than you think: the role of categories in visual long-term memory. Psychol. Sci.21, 1551–1556.10.1177/0956797610385359
37
LammeV. A. F.RoelfsemaP. R. (2000). The distinct modes of vision offered by feedforward and recurrent processing. Trends Neurosci.23, 571–579.10.1016/S0166-2236(00)01657-X
38
LiuH.AgamY.MadsenJ. R.KreimanG. (2009). Timing, timing, timing: fast decoding of object information from intracranial field potentials in human visual cortex. Neuron62, 281–290.10.1016/j.neuron.2009.02.025
39
LoftusG. R.GinnM. (1984). Perceptual and conceptual masking of pictures. J. Exp. Psychol. Learn. Mem. Cogn.10, 435–441.10.1037/0278-7393.10.3.435
40
LoftusG. R.HannaA. M.LesterL. (1988). Conceptual masking: how one picture captures attention from another picture. Cogn. Psychol.20, 237–282.10.1016/0010-0285(88)90020-5
41
NickersonR. S. (1965). Short-term memory for complex meaningful visual configurations: a demonstration of capacity. Can. J. Psychol.19, 155–160.10.1037/h0082899
42
OlivaA. (2005). “Gist of the scene,” in Encyclopedia of Neurobiology of Attention, eds IttiL.ReesG.TsotsosJ. K. (San Diego, CA: Elsevier), 251–256.
43
PerrettD.HietanenJ.OramM.BensonP. (1992). Organization and functions of cells responsive to faces in the temporal cortex. Philos. Trans. R. Soc. Lond. B Biol. Sci.335, 23–30.10.1098/rstb.1992.0003
44
PhillipsW. A. (1983). Short-term visual memory. Philos. Trans. R. Soc. Lond. B Biol. Sci302, 295–309.10.1098/rstb.1983.0056
45
PhillipsW. A.ChristieD. F. M. (1977). Components of visual memory. Q. J. Exp. Psychol. (Hove)29, 117–133.
46
PotterM. C. (1975). Meaning in visual search. Science187, 965–966.10.1126/science.1145183
47
PotterM. C. (1976). Short-term conceptual memory for pictures. J. Exp. Psychol. Hum. Learn. Mem.2, 509–522.10.1037/0278-7393.2.5.509
48
PotterM. C. (1983). “Representational buffers: the eye-mind hypothesis in picture perception, reading, and visual search,” in Eye Movements in Reading: Perceptual and Language Processes, ed. RaynerK. (New York: Academic Press), 423–437.
49
PotterM. C. (1993). Very short-term conceptual memory. Mem. Cognit.21, 156–161.10.3758/BF03202727
50
PotterM. C. (1999). “Understanding sentences and scenes: the role of conceptual short term memory,” in Fleeting Memories: Cognition of Brief Visual Stimuli, ed. ColtheartV. (Cambridge, MA: MIT Press), 13–46.
51
PotterM. C. (2010). Conceptual short term memory. Scholarpedia5, 333410.4249/scholarpedia.3334
52
PotterM. C.FaulconerB. A. (1975). Time to understand pictures and words. Nature253, 437–438.10.1038/253437a0
53
PotterM. C.FoxL. F. (2009). Detecting and remembering simultaneous pictures in a rapid serial visual presentation. J. Exp. Psychol. Hum. Percept. Perform.35, 28–38.10.1037/a0013624
54
PotterM. C.JiangY. V. (2009). “Visual short-term memory,” in Oxford Companion to Consciousness, eds BayneT.CleeremansA.WilkenP. (Oxford: Oxford University Press), 436–438.
55
PotterM. C.LevyE. I. (1969). Recognition memory for a rapid sequence of pictures. J. Exp. Psychol.81, 10–15.10.1037/h0027470
56
PotterM. C.MoryadasA.AbramsI.NoelA. (1993). Word perception and misperception in context. J. Exp. Psychol. Learn. Mem. Cogn.19, 3–22.10.1037/0278-7393.19.1.3
57
PotterM. C.StaubA.O’ConnorD. H. (2004). Pictorial and conceptual representation of glimpsed pictures. J. Exp. Psychol. Hum. Percept. Perform.30, 478–489.10.1037/0096-1523.30.3.478
58
PotterM. C.StaubA.RadoJ.O’ConnorD. H. (2002). Recognition memory for briefly-presented pictures: the time course of rapid forgetting. J. Exp. Psychol. Hum. Percept. Perform.28, 1163–1175.10.1037/0096-1523.28.5.1163
59
PotterM. C.StiefboldD.MoryadasA. (1998). Word selection in reading sentences: preceding versus following contexts. J. Exp. Psychol. Learn. Mem. Cogn.24, 68–100.10.1037/0278-7393.24.1.68
60
PotterM. C.WybleB.PandavR.OlejarczykJ. (2010). Picture detection in RSVP: features or identity?J. Exp. Psychol. Hum. Percept. Perform.36, 1486–1494.10.1037/a0018730
61
RensinkR. A.O’ReganJ. K.ClarkJ. J. (2000). On the failure to detect changes in scenes across brief interruptions. Vis. Cogn.7, 127–145.10.1080/135062800394847
62
RensinkR. A.O’ReganJ. R.ClarkJ. J. (1997). To see or not to see: the need for attention to perceive changes in scenes. Psychol. Sci.8, 368–373.10.1111/j.1467-9280.1997.tb00427.x
63
RousseletG.Fabre-ThorpeM.ThorpeS. J. (2002). Parallel processing in high level categorization of natural images. Nat. Neurosci.5, 629–630.
64
RousseletG. A.ThorpeS. J.Fabre-ThorpeM. (2004a). Processing of one, two or four natural scenes in humans: the limits of parallelism. Vision Res.44, 877–894.10.1016/j.visres.2003.11.014
65
RousseletG.ThorpeS. J.Fabre-ThorpeM. (2004b). How parallel is visual processing in the ventral pathway?Trends Cogn. Sci. (Regul. Ed.)8, 363–370.10.1016/j.tics.2004.06.003
66
SerreT.KreimanG.KouhM.CadieuC.KnoblichU.PoggioT. (2007a). A quantitative theory of immediate visual recognition. Prog. Brain Res.165, 33–56.10.1016/S0079-6123(06)65004-8
67
SerreT.OlivaA.PoggioT. (2007b). A feedforward architecture accounts for rapid categorization. Proc. Natl. Acad. Sci. U.S.A.104, 6424–6429.10.1073/pnas.0700622104
68
ShepardR. N. (1967) , Recognition memory for words, sentences, and pictures. J. Mem. Lang.6, 156–163.
69
SimonsD. J.LevinD. T. (1997). Change blindness. Trends Cogn. Sci. (Regul. Ed.)1, 261–267.10.1016/S1364-6613(97)01080-2
70
SperlingG. (1960). The information available in brief visual presentations. Psychol. Monogr.74, 1–29.10.1037/h0093761
71
StandingL. (1973). Learning 10,000 pictures. Q. J. Exp. Psychol. (Hove)25, 207–222.
72
ThorpeS.Fabre-ThorpeM. (2001). Seeking categories in the brain. Science291, 260–263.10.1126/science.1058249
Summary
Keywords
picture perception, rapid serial visual presentation, picture memory, detection, feedforward processing, masking, search
Citation
Potter MC (2012) Recognition and Memory for Briefly Presented Scenes. Front. Psychology 3:32. doi: 10.3389/fpsyg.2012.00032
Received
30 November 2011
Accepted
28 January 2012
Published
22 February 2012
Volume
3 - 2012
Edited by
Gabriel Kreiman, Harvard Medical School, USA
Reviewed by
Rufin VanRullen, Centre de Recherche Cerveau et Cognition Toulouse, France; Gabriel Kreiman, Harvard Medical School, USA
Copyright
© 2012 Potter.
This is an open-access article distributed under the terms of the Creative Commons Attribution Non Commercial License, which permits non-commercial use, distribution, and reproduction in other forums, provided the original authors and source are credited.
*Correspondence: Mary C. Potter, Department of Brain and Cognitive Sciences, Massachusetts Institute of Technology, 46-4125, 77 Massachusetts Avenue, Cambridge, MA 02139, USA. e-mail: mpotter@mit.edu
This article was submitted to Frontiers in Perception Science, a specialty of Frontiers in Psychology.
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.