Abstract
For adaptive real-time behavior in real-world contexts, the brain needs to allow past information over multiple timescales to influence current processing for making choices that create the best outcome as a person goes about making choices in their everyday life. The neuroeconomics literature on value-based decision-making has formalized such choice through reinforcement learning models for two extreme strategies. These strategies are model-free (MF), which is an automatic, stimulus–response type of action, and model-based (MB), which bases choice on cognitive representations of the world and causal inference on environment-behavior structure. The emphasis of examining the neural substrates of value-based decision making has been on the striatum and prefrontal regions, especially with regards to the “here and now” decision-making. Yet, such a dichotomy does not embrace all the dynamic complexity involved. In addition, despite robust research on the role of the hippocampus in memory and spatial learning, its contribution to value-based decision making is just starting to be explored. This paper aims to better appreciate the role of the hippocampus in decision-making and advance the successor representation (SR) as a candidate mechanism for encoding state representations in the hippocampus, separate from reward representations. To this end, we review research that relates hippocampal sequences to SR models showing that the implementation of such sequences in reinforcement learning agents improves their performance. This also enables the agents to perform multiscale temporal processing in a biologically plausible manner. Altogether, we articulate a framework to advance current striatal and prefrontal-focused decision making to better account for multiscale mechanisms underlying various real-world time-related concepts such as the self that cumulates over a person’s life course.
1. Introduction
After a long day at work, it is time to go home. If one has worked in the same building for several years, one does not actively think about how to get out of the building as a key milestone on the way to the goal of getting home. This action simply involves coming out of the elevator and turning left or right to exit onto the street. This simple decision can be a bit different, though, if construction in the building blocks the exit. Instead of getting out of the elevator as usual, one may remember a nearby fire exit and get out from the building.
This extremely simplistic example of a real-world behavior sequence in reaching a goal has been used in previous psychological research in human decision-making (; Wood et al., 2021). Clearly, making a choice with the best outcome is more complex than this simplistic example most of the time. It is also a lifelong challenge that requires considering outcomes on different timescales and calling for adaptation to stable, diverse, and changing contexts that one encounters every day and over one’s life course ().
Examining adaptive behavior from such a lifespan behavioral and decision neuroscience perspective calls upon the interface of value-based decision-making with literature on learning, memory, and spatial navigation (). The literature on value-based decision making has formalized two types of reinforcement learning strategies used in decision-making: model-free (MF), which is an automatic, stimulus–response type of action. The other type of strategy is called model-based (MB), wherein we use the knowledge of the cognitive representation of the world around us and causal environment-behavior inference to plan our next action with more flexibility but also less efficiency. MB action is an important component of planning and deliberative decision-making, where one needs to mentally imagine future scenarios and make a choice, something which we routinely do in our lives. MB ability has shown a developmental pattern, progressively emerging with age from childhood to adulthood ().
Broadly speaking, the MF and MB strategies have been attributed to distinct brain regions; the striatum is involved in automatic responses which are the hallmark of MF strategies, while the hippocampus is thought to be key not only for episodic and spatial memories but also for building a model of the world.
The striatum is associated with the dopaminergic system and reward, as well as neural representations that track value. These properties have resulted in the striatum becoming a hotbed of focus for researchers studying neuroeconomics, decision-making and behavioral neuroscience. Fundamental studies in psychology (Pavlov, 1960; ) have lent themselves well to quantitative approaches, facilitating the growth and emergence of several computational models of reinforcement learning (See Samson et al., 2010 for review). Computational models of deliberation and planning in the brain are relatively more recent (Mattar and Lengyel, 2022; ), and consequently the interactions between the hippocampal and striatal system have only recently garnered attention.
For example, showed that the hippocampus supports both spatial and habitual memories when these events have temporal proximity, while the striatum supports both types of memories for events sharing a common spatial context. Models unifying the two systems have also been presented in relation to decision-making. introduced a model in which the hippocampal-striatal system was viewed as a general system for decision making via an adaptive combination of the MF and MB frameworks. However, more research is needed to better understand the dynamics of interaction between the two systems, especially in real-world contexts, which are ever-changing ().
Taken together, there is a need to move beyond the present false dichotomy implying that value-based decision making is either MF or MB and better appreciate the role of the hippocampus in decision-making, beyond its role in the encoding of episodic memories (Scoville and Milner, 1957). This is exemplified by RL models using replay as a strategy to improve task performance (Russek et al., 2017; van de Ven et al., 2020). The hippocampus too exhibits replay, suggesting that it could be contributing to reward learning in the brain. Understanding this crucial link opens avenues to understanding hippocampal contributions to decisions in real-world timescales as well as long-term decisions as episodes accumulate over a person’s life course, impacting an individual’s health and overall wellness.
In the subsequent sections, we will review RL models and how they inform our thinking about neural processes. In relation to these models, we will introduce the hippocampus as a sequence generator (). We will review specific examples of hippocampal sequences to demonstrate that these sequences can be used for MB actions. While most of this work has been done in the rodent spatial navigation system, the prevailing notion is that these sequences are attributed with meaningful content as the animal experiences its environment (). Thus, hippocampal sequences can be generalized to any form of multimodal information that is sequential in nature. Studying spatial navigation simply provides a convenient, tractable foray into understanding hippocampal function. Specifically, we will argue that the hippocampal provides key neural substrates for the continuous sense of the self over time along a person’s life course. Finally, we will suggest potential ideas for interdisciplinary convergence bridging animal systems to human neuroscience studies as well as to the fields of neuroeconomics.
2. RL models of decision-making
The main aim of RL models is for an agent to maximize its reward given a state and learn the optimal action policy for it to do so. For example, these states could be locations, stimuli, or reward contingencies. RL agents broadly fall into two categories: MF and MB. MF agents learn via prediction errors between the expected value of the state and the observed value, a process known as temporal difference (TD) learning. These agents try to minimize the prediction error between observed and expected reward value over the long term. This approach is typically computationally faster and cheaper, but such agents have no memory of past states or the relationship between them. For example, if the value of the reward changes, a MF agent will only be able to update itself by revisiting different states several times and experiencing the consequences of its actions repeatedly.
On the other hand, MB agents learn a representation of the various states and transitions between them and use this to maximize long-term expected reward value. Representing the entire set of transitions and states is what makes these models flexible to novel task demands. Unlike a MF agent, a MB agent would not need to experience states repeatedly to learn changes in reward values; the process is much faster since the agent is able to exploit the task structure to learn such changes. However, MB algorithms are significantly more computationally intensive.
RL approaches lend themselves well to biology and provide a framework to generate testable predictions about the working of the striatal system in the context of reward learning and decision-making. In the brain, the striatum is thought to signal value, which is updated by a dopaminergic prediction error signal (Schultz et al., 1997), and these dopamine responses are in line with predictions of TD learning models (Waelti et al., 2001).
How the brain learns the model of the environment is a relatively more complex question to tackle. Unlike MF approaches which emphasize stimulus–reward associations, animals do not even need reward to learn the structure of the environment. The retention of information in the absence of external reinforcers is referred to as latent learning (Tolman, 1948). Latent learning enables the animal to quickly predict future rewards, or even generalize learnt knowledge to other state spaces. Such mental representations of the environment came to be known as cognitive maps. How cognitive maps contribute to MB agents remains a gap in the field. In the 1970s, the discovery of place cells (O’Keefe, 1976) brought the hippocampus into focus as the seat of the cognitive map. More recent studies in humans have begun to show the direct contributions of the hippocampus to MB planning (Miller et al., 2017; Vikbladh et al., 2019). Therefore, understanding hippocampal function is an essential avenue to further our knowledge of MB decision-making.
3. The successor representation
Current experimental setups sometimes fail to accurately assess how and to what extent an agent uses MF and MB strategies in decision-making problems, and a re-evaluation of the assumptions underlying these strategies is much needed (). This would then allow for further investigation into the role of the hippocampus within MF and MB actions more clearly.
Therefore, more recent methodologies in RL emphasize a combination of MF and MB agents to improve the generalization of TD learning approaches. One such approach that has gained popularity in neuroscience is the successor representation (SR) (; ).
RL provides a formal means of investigating decision-making, in which states of the world are rewarded and decisions must be made on the selection of actions that can be taken to maximize reward. In this framework, each state has a value (V), defined as the cumulative expected reward over future states, multiplied by a discount factor () that reduces the weight of distal rewards.
It is useful to make decisions based on the estimated value of different states. As shown by , the value function can be mathematically represented as the inner product of the reward function (R) and a representation of the estimated value of states (M), as shown below:
The matrix M is the SR matrix. The SR possesses a state representation which conveys the discounted number of expected visits of a given future state (s’) from a given starting state (s). The SR matrix is given by:
Where T is the transition matrix, and t denotes all future time steps in the planning horizon. Instead of computing the transition matrix for each step, the SR is computed as a discounted sum going from state s to state s’ in a given number of steps, determined by the planning horizon (Figure 1A). This representation therefore has predictive structure, akin to a MB agent, but can be learned by a MF agent via TD learning, by learning the difference between observed and expected state occupancy. The SR approach thereby integrates the advantages of a MB agent into a MF framework (Figure 1B).
Figure 1
From Equation (1), we observe that the SR is a representation of possible future states that can be separable from the value function. In a reward revaluation task (i.e., change in reward value), this allows the agent to retain the same predictive map and quickly compute the value, whereas an MB agent would have to recompute the mapping between states, and an MF agent would have to re-learn the environment altogether. The SR thus offers an optimal solution to this kind of task and permits the learning of the state transitions (or “map”) independently of reward.
Additionally, Equation (1) also provides a direct relationship between the value function and SR, suggesting that updates to SR can update the value function. The SR, therefore, forms an important link between predictive representations and the value-based decision-making framework.
The SR, however, has its own set of caveats. It requires direct experience to learn, akin to an MF agent. If the transition structure between states were to change through the course of the task (known as transition revaluation), the SR would only be able to update the one-step transition but not the steps preceding this state (Figure 1C). This is because the SR is probabilistic and has no temporal representation built into it. As a result, SR agents are unable to solve transition revaluation or policy revaluation (i.e., change in strategy) tasks, which animals can easily adapt to Tolman (1948) and Simon and Daw (2011).
Despite the caveats, SR models have been of increasing interest in neuroscience due to their biological plausibility, accompanied with observations of their behavioral and neural correlates during decision-making. Using a sequential learning task (Momennejad et al., 2017), human participants learnt a relationship between stimulus and reward, which was manipulated in the re-learning phase, and subsequently probed in the final phase of the task. In the re-learning phase, the investigators performed either a reward revaluation, or a transition revaluation. As detailed above, an SR agent would be able to solve reward revaluation but not transition revaluation. This was recapitulated with the participants; they were able to adjust better to reward revaluation compared to transition revaluation, suggesting the utilization of cached representations (analogous to SR) to solve the task.
In addition, the SR has neural correlates in the hippocampus: If a SR agent is allowed to forage in an open arena with uniformly distributed rewards and there exists a population of neurons encoding each spatial state, the neural population activity (i.e., the columns of the SR matrix) resembles hippocampal place fields (spatial locations where cells fire most, in an arena) (Stachenfeld et al., 2017). In the same study, the authors showed that the eigenvectors of the SR matrix resemble grid cells (cells that fire in a hexagonal grid-like pattern within a given environment). Predictions from such models also recapitulated experimental observations such as the clustering of place fields around rewarded locations (
Recent work on the SR has tried to address and resolve the lack of temporal resolution in the SR. Momennejad and Howard (2018) showed that an ensemble of SR matrices with different discount factors (denoting different timescales) can be used to incorporate sequential order by encoding the Laplace transform of the future. A Laplace transform decomposes a signal into exponential decay functions of different rates. The inverse of this is equivalent to computing a derivative of the relation between two given states across SR matrices, i.e., across timescales. This consequently enables recovery of the temporal order between states. The mathematical formulation of this approach resembles that used in the temporal context model, detailed in a later section (See section: A broader view of hippocampal sequences).
A prediction that arises from multi-scale SR is the presence of cells that are sequentially activated as a function of the distance to the goal (Momennejad and Howard, 2018). Such cells have been experimentally observed in the hippocampus of bats and mice (Sarel et al., 2017;
Additional support for multi-scale SR in the brain, comes from a study by
Interestingly, predictive horizons analyzed in
Taken together, these observations provide evidence for the utility of SR in using RL-based approaches to understand the neural representations of space in the brain and more directly exhibit the predictive nature of hippocampal representations (Stachenfeld et al., 2017), suggesting that a multi-scale SR might be implemented across brain regions, spanning the hippocampus to the PFC, warranting further investigation into the mechanisms behind how these regions communicate during real-world decisions.
In summary, the SR is an RL-based framework of predictive representations that combines some of the speed of MF and the flexibility of MB agents. Such a predictive system is reminiscent of the hippocampal memory system, as evidenced from various neural correlates of the SR in the hippocampus. Most models of hippocampal function focus on learning (Uria et al., 2020; Whittington et al., 2020;
4. Linking SR models to hippocampal sequences
The SR being a state-based model relies on the delineation of explicit states that the agent can be in at any given time. In a computational agent, these states are explicitly encoded. However, if animals were to implement the SR, these states are likely learned and updated from experience. The learning of the SR, therefore, is an interesting research direction that warrants future work that can potentially inform real-world decision-making.
Traditionally, SR models were learnt using TD learning, which is not known to be implemented in the hippocampal circuitry (but see
Work on how the SR is learnt and updated has also given rise to models that perform better than classical SR models and provide not only a better understanding of hippocampal function, but also lend valuable insights into real-time decision-making in real-world contexts. Russek et al. (2017) introduced an SR agent that can solve transition and policy revaluation tasks, called SR-Dyna. This agent learns representations through online experience, and in addition prioritizes recent experience using “offline replay,” referring to the simulation of experiences by playing back past episodes (
Offline replay has been shown to be a key process in contributing to generalization and memory consolidation in several human and animal studies (
The existence of anatomical substrates for the integration of replay into the reward learning system (thought to be implemented by the striatum) makes SR-Dyna well-poised to further understand the role of hippocampus in decision-making. In particular, the hippocampus and the dopamine system form an anatomical loop; the hippocampus receives dopaminergic inputs from the ventral tegmental area (VTA) and in turn projects to the ventral striatum (nucleus accumbens) and globus pallidus, which projects back to the VTA (
In summary, RL approaches to decision-making have provided a mathematical framework to understand how the brain can possibly implement reward learning, and thereby pursue the strategy that leads to maximal expected reward (Samson et al., 2010). However, we still lack a comprehensive understanding of how the brain learns the structure of the environment and implements MB algorithms for efficient decision-making. We propose that this gap in the field can be bridged by appreciating the role of the hippocampus in decision-making.
In subsequent sections, we will review further evidence for the role of the hippocampus in prospective coding, i.e., future planning of actions. To do so, we will further build upon hippocampal replay, which is a form of offline consolidation. We will then introduce a substrate for prospective coding known as theta sequences (
5. Linking the successor representation to the memory system
In the 1940s, Tolman (1948) performed behavioral experiments with rats navigating a maze. Once the rats learnt the reward location, the shape of the maze was drastically altered. Yet, the rats were able to efficiently navigate to the same reward location, regardless of the shape of the maze. This observation suggested that the animals were able to form a representation of spatial location, without any direct stimulus–reward association, known as latent learning. This also begged the question of the neural correlates of the cognitive map that the animals used to reach the reward, long before the arrival of computational models of RL.
Fast-forward to the 1970s. With advances in electrophysiology, it became possible to record neurons in freely moving animals. This led to John O’Keefe’s discovery of place cells in the hippocampus (O’Keefe, 1976). This discovery generated interest in the hippocampus as the seat of the cognitive map. Subsequently, it was found that the hippocampus represents several other neural representations based on what is salient information for the task at hand, such as time (Pastalkova et al., 2008), a conspecific (
In summary, the prevailing theory of hippocampal function is thought to be the binding of spatial, temporal, and other sensory features into an episode, thereby being important for episodic memories (
6. Hippocampal replay
Once it became possible to simultaneously record several neurons in the hippocampus, researchers could now investigate the population activity of the hippocampus. Recording several place cells as a rat ran around in a freely moving arena, Wilson and McNaughton (1994) observed that neurons that tended to fire together when the animal was exploring an arena also tend to fire together during post-task sleep. Such reactivations usually occur during non-Rapid Eye Movement (NREM) sleep or during quiet wakefulness when the animal is disengaged from its environment, such as during grooming or consummatory behaviors (
Reactivation events that have a temporal sequence (for example, a sequence that corresponds to a trajectory of place fields) are said to be replayed (Figure 2A). Hippocampal replay and offline replay as implemented in SR-Dyna have direct parallels, since both involve the recapitulation of previously experienced events and states, respectively. Importantly, hippocampal replay is thought to serve as a substrate for memory consolidation. Impairing SWRs or prolonging them can worsen or improve task performance, respectively (
Figure 2

Overview of spatial sequences in the hippocampus. (A) Hippocampal replay: (Left) Firing of hippocampal place cells as a rat runs on a maze. The cells are successively activated as the animal traverses through their place fields (color-coded on the maze), forming a sequence. (Center) Place cells indicate spatial location by firing maximally at their preferred location (known as a place field). (Right) Sharp-wave ripples (SWRs) occur during NREM sleep or quiet wakefulness and is associated with increased hippocampal population activity. During a SWR, place cell trajectories that were experienced during wakefulness are “replayed.” Adapted with permission from Zielinski et al. (2017). (B) Encoding of spatial location within theta sequences. (Left) The location of the animal is encoded via a phase code of the theta oscillation, with past locations being represented on the negative phase and future locations on the positive phase of the theta oscillation. The current location is represented at the trough of the oscillation. Reprinted from Petersen and Buzsáki (2020), with permission from Elsevier. (Right) During a deliberative decision task, hippocampal population activity nested within theta sequences sweeps forward in time, representing future spatial options. Reproduced from Redish (2016) with permission from SNCSC.
7. Theta sequences and prospective coding in the hippocampus
In addition to SWRs, which are a form of offline consolidation, the hippocampus also exhibits sequences during online planning. These sequences are known as theta sequences, named after the theta oscillation, a characteristic oscillation of population activity between 8–12 Hz which is observed during locomotion or during Rapid Eye Movement (REM) sleep. Theta sequences can provide important insights into understanding the here-and-now type of deliberative decision-making in real-world scenarios.
O’Keefe and Recce (1993) discovered that as animals traversed across a linear track, the spiking activity of place cells shifted to earlier phases of the ongoing theta oscillation. This phenomenon is known as theta phase precession and is a crucial component for the encoding of place cell sequences (Skaggs et al., 1996; O’Keefe and Burgess, 2005) (Figure 2B, left). More recently, theta sequences were shown to be directly implicated in prospective coding.
Additional evidence of the involvement of the hippocampus in prospective coding comes from a study done by Ito et al. (2015). They examined the activity of hippocampal neurons that fire differently based on the animal’s past or future behavior, known as splitter cells (
In line with the role of theta sequences in prospective coding,
8. A broader view of hippocampal sequences
Hippocampal sequences are thought to be essential features for encoding episodic memory. This is because episodic memory is also sequential in nature. However, along with spatial details, these episodes also typically involve temporal details. In particular, understanding the neural basis of timing is important to understand memory-guided decision-making, because when we make decisions, we typically recall events that may go back to several years ago. Below, we will touch upon temporal sequences in the hippocampus and how an understanding of these sequences can inform decision-making research.
Studies by
The discovery of time cells in the hippocampus (Pastalkova et al., 2008) showed that the hippocampus can use sequences to encode time intervals leading up to the end of a delay period. Along with this, other findings showing the evolution of hippocampal activity over hour-long intervals (Manns et al., 2007) suggest that the hippocampus utilizes sequences to represent different time scales as well.
One prevailing theory for the representation of time is known as the temporal context model (TCM) (
In summary, spatiotemporal sequences are thought to be the neural substrates for human episodic memory that is vivid and rich with spatiotemporal and other multimodal information such as olfaction; the smell of our mother’s cooking can take us back to our childhood in a flash. The hippocampus is thus thought to bind all these features together into a coherent representation of memory. Understanding how the reward learning system utilizes the information encoded in hippocampal sequences representing such a vast diversity of information to guide decisions is an exciting direction of research for the decision-making field.
9. Beyond the rodent hippocampus
Extending the findings from rodent studies to humans is important to understand the mechanisms of decision-making. However, the techniques used to study the precise timing of hippocampal sequences are difficult to directly be applied to human research for various reasons. First, electrophysiological recordings are invasive, and are therefore only performed on patients who are being monitored for surgical removal of epileptic tissue. These patients often have altered brain activity and impaired decision-making, making it difficult to study what happens in a healthy individual. Having access to a good sample size of patients is an additional challenge. Second, non-invasive techniques such as fMRI are useful to study healthy individuals, but have poorer temporal resolution, making it a challenge to study hippocampal sequences such as SWRs and theta sequences which occur on the scale of milliseconds.
Despite these challenges, research in humans is catching up with the advances made in rodent spatial navigation with the demonstrations of place, grid, and time cells, using single unit recordings and fMRI (
Recent research with human subjects offers promising avenues for the role of the hippocampus in learning and decision-related activity. Using fMRI in infants,
Despite being limited by measuring vascular responses and poor spatiotemporal resolution, fMRI offers whole-brain access, which is not as easy in rodents with current techniques. This has led to deeper insights on how the hippocampus, in conjunction with the prefrontal cortex and striatum, represents abstract information during decision-making tasks, such as representations of task structure from experience in conjunction with the orbitofrontal cortex (Mızrak et al., 2021), the combination of spatial and non-spatial variables during goal-directed decision-making (Viard et al., 2011), and deliberation during value-based decision making (
Such whole-brain studies have also led to the characterization of other distinct brain network modules, such as the dorsal and ventral attention networks, the default mode network, and the visual network to name a few (Power et al., 2010).
These brain networks have confirmed that regions that were thought to work in synchrony are indeed co-modulated during tasks such as attention, memory, and decision-making. Notably, the hippocampus along with the prefrontal cortex is part of the default mode network (DMN), a network thought to be active when we are not engaged in any task but are introspecting, deliberating, or recalling past experiences (
10. Hippocampal contributions to understanding the self and lifelong real-world decision making
The field of neuroeconomics is predominantly inundated with research in value-based decision-making, focusing on the striatum and prefrontal cortex. As the self embodies a person’s lifelong decision-making and experience, can a deeper and more comprehensive mapping of the interaction of hippocampus with the striatum and prefrontal regions open new horizons for scaling up the impact of neuroscience research on real-world applications for better mental health and wellbeing?
The self is at the core of our mental life, creating a continuous thread guiding decision-making over the course of a person’s lifespan and as a function of real-time and cumulative experience and context (
The multi-functional view of the self featured above provides important insights on how the double integration of the self and hippocampus in current neuroeconomics approaches to value-based choice can advance a real-world decision-making framework that is biological, culturally and psychologically plausible. As mentioned at the onset, the prevailing assumption in decision neuroscience and neuroeconomics is that the brain reward system encodes representations of the online expected value of stimuli and/or actions through the ventromedial prefrontal cortex (vmPFC), supplementary motor area, and the striatum (
The field of neuroscience has started to explore the neural mechanisms underlying our sense of self (see
We now know that accounts of value-based decision-making, (and potentially also of the self), are incomplete without the incorporation of the hippocampus and would therefore like to usher a change in the framework of current decision-making research by advocating that the hippocampus is a crucially essential component underlying decision-making and the self.
As discussed in this review, the hippocampus can potentially implement a predictive framework of the environment, exemplified by the neural correlates of SR models. As shown by multi-scale SR models, this predictive representation can span multiple timescales. Understanding how memory and time are integrated in the brain offers an attractive avenue for understanding the relationship between the episodic memory system, the self, and decision-making, mediated by the hippocampus and the DMN, in conjunction with the striatal reward learning system. In addition to the hippocampus, interval timing is also represented in the prefrontal and parietal cortical areas and is thought to be integrated via the striatum (
Finally, recent neuroimaging evidence sheds light on non-value signals from the hippocampus and the rest of the DMN (
11. Discussion
Due to the multi-disciplinary nature of the decision-making field, several attractive directions of research present themselves in the concluding section of this article, having far-reaching implications not only in the fields of computational modeling, but also spatial navigation, neuroeconomics, marketing, and health. Advances in these disciplines will, in turn, provide an integrative account of the mechanisms underlying real-world decision-making along a person’s life course.
We now have powerful computational tools capable of solving a variety of “here-and-now” decision-making tasks and are beginning to understand the mechanisms behind credit assignment of future states, which evidence is important during credit assignment how belief updating is integrated into the memory system, and how this leads to changes in policy. The hippocampus is poised to serve crucial roles in these processes, making it relevant to decision-making researchers.
One exciting avenue in accounting for the role of hippocampal replay in decision-making over time lies in the domain of continual learning (CL). In CL, a model must continually learn new tasks, while maintaining its performance on previously learnt tasks. A major challenge in CL is catastrophic forgetting, in which performance on previously acquired tasks drops due to interference with learning of a new task. However, we know that animals can learn many different concepts throughout their lifetime without forgetting previously learnt ideas. Interestingly, using replay to counteract catastrophic forgetting is emerging as a popular approach in the field (van de Ven et al., 2020;
However, the role of theta sequences in a CL agent is less clear. SWRs and theta sequences co-exist in the hippocampus with an inverse relationship; ripples occur during so-called offline states, while theta sequences occur during real-time decision-making. Novel computational approaches can provide answers to which process is important when and for what kind of tasks. Specifically, disambiguating the exact roles that these oscillations play in the context of learning, consolidation and planning will directly be of benefit for biologically plausible modeling of neural representations and behavior. In addition, understanding the content and valence associated with hippocampal sequences and how these factors are integrated with the striatal and prefrontal systems will bring more clarity in our understanding of the meaning behind the activation of brain networks such as the DMN post-ripple.
CL models using generative replay approaches may also contribute to an understanding of autobiographical memories. By continuously encoding and replaying episodes, such agents may provide an account of events that persists through time. By combining the past with the present and prospecting about the future, this can advance the understanding of the neural correlates of the sense of self and its roles in episode-specific and lifelong learning and reward (
12. Conclusion
The overarching message of this article is to portray the hippocampus as a key but understudied aspect of human decision-making. By using past information in the form of memory to guide our future actions, the hippocampus could well be a core neural substrate of decision-making per se, and its interaction with the striatum and prefrontal regions – continuously updating and modulating the MF-MB balance, thereby impacting real-world decision-making in the “here and now” as well as on the long term, with the accumulation of experiences and contexts as the self, unfolding over time and across scales and dimensions (Northoff and Hayes, 2011;
Our species has reached its current state through the evolution of a highly sophisticated brain engaging in decision-making that ranges from canonical MF and MB to something in between, with the self being one of the most distinguishing facets of human evolution. This enables real-world behavior to be adaptive to an ever more complex and dynamic immediate environment as well as to social institutions and globe-spanning digital communities. Creating a world that supports multiscale computational efficiency and resilience in human and machine is a pressing necessity (
Funding
DM is funded by a doctoral training award by the Fonds de Recherche du Québec, Santé. LD is funded by the Driving adaptive versus materialistic consumption to benefit consumers and marketers. Grant # SSHRC 435-2020-1136. Implementing Smart Cities Interventions to Build Healthy Cities. Healthy Cities Training. Grant # CIHR-NSERC-SSHRC/Guelph 02083-000.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Statements
Author contributions
DM performed the literature review and wrote the first draft of the manuscript. DM and LD wrote sections of the manuscript. All authors contributed to the article and approved the submitted version.
Acknowledgments
The authors would like to thank Dr. Gina Kemp for proofreading an earlier version of the manuscript, Göktuğ Bender, and Alexandra Paquette for assistance with manuscript revision, and Dr. Daniel Levenstein for helpful discussions.
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
References
1
AbadchiJ. K.Nazari-AhangarkolaeeM.GattasS.Bermudez-ContrerasE.LuczakA.McNaughtonB. L.et al. (2020). Spatiotemporal patterns of neocortical activity around hippocampal sharp-wave ripples. elife9, 1–26. doi: 10.7554/eLife.51972
2
AddisD. R.ChengT. P.RobertsR.SchacterD. L. (2011). Hippocampal contributions to the episodic simulation of specific and general future events. Hippocampus21, 1045–1052. doi: 10.1002/hipo.20870
3
AlvernheA.SaveE.PoucetB. (2011). Local remapping of place cell firing in the Tolman detour task. Eur. J. Neurosci.33, 1696–1705. doi: 10.1111/j.1460-9568.2011.07653.x
4
AquinoT. G.CockburnJ.MamelakA. N.RutishauserU.O’DohertyJ. P. (2023). Neurons in human pre-supplementary motor area encode key computations for value-based choice. Nat. Hum. Behav.7, 970–985. doi: 10.1038/s41562-023-01548-2
5
AronovD.NeversR.TankD. W. (2017). Mapping of a non-spatial dimension by the hippocampal-entorhinal circuit. Nature543, 719–722. doi: 10.1038/nature21692
6
AxmacherN.ElgerC. E.FellJ. (2008). Ripples in the medial temporal lobe are relevant for human memory consolidation. Brain131, 1806–1817. doi: 10.1093/brain/awn103
7
BackusA. R.SchoffelenJ. M.SzebényiS.HanslmayrS.DoellerC. F. (2016). Hippocampal-prefrontal theta oscillations support memory integration. Curr. Biol.26, 450–457. doi: 10.1016/j.cub.2015.12.048
8
BakkourA.PalomboD. J.ZylberbergA.KangY. H. R.ReidA.VerfaellieM.et al. (2019). The hippocampus supports deliberation during value-based decisions. elife8, 1–28. doi: 10.7554/eLife.46080
9
BalleineB. W.DelgadoM. R.HikosakaO. (2007). The role of the dorsal striatum in reward and decision-making. J. Neurosci.27, 8161–8165. doi: 10.1523/JNEUROSCI.1554-07.2007
10
BenchenaneK.PeyracheA.KhamassiM.TierneyP. L.GioanniY.BattagliaF. P.et al. (2010). Coherent Theta oscillations and reorganization of spike timing in the hippocampal- prefrontal network upon learning. Neuron66, 921–936. doi: 10.1016/j.neuron.2010.05.013
11
BidermanN.BakkourA.ShohamyD. (2020). What are memories for? The hippocampus bridges past experience with future decisions. Trends Cogn. Sci.24, 542–556. doi: 10.1016/j.tics.2020.04.004
12
BonoJ.ZannoneS.PedrosaV.ClopathC. (2023). Learning predictive cognitive maps with spiking neurons during behaviour and replays. elife12:e80671. doi: 10.7554/eLife.80671
13
BrunecI. K.MomennejadI. (2022). Predictive representations in hippocampal and prefrontal hierarchies. J. Neurosci.42, 299–312. doi: 10.1523/JNEUROSCI.1327-21.2021
14
BucknerR. L.CarrollD. C. (2007). Self-projection and the brain. Trends Cogn. Sci.11, 49–57. doi: 10.1016/j.tics.2006.11.004
15
BuzsákiG.TingleyD. (2018). Space and time: the hippocampus as a sequence generator. Trends Cogn. Sci.22, 853–869. doi: 10.1016/j.tics.2018.07.006
16
CarrM. F.JadhavS. P.FrankL. M. (2011). Hippocampal replay in the awake state: a potential substrate for memory consolidation and retrieval. Nat. Neurosci.14, 147–153. doi: 10.1038/nn.2732
17
DanjoT.ToyoizumiT.FujisawaS. (2018). Spatial representations of self and other in the hippocampus. Science359, 213–218. doi: 10.1126/science.aao3898
18
DayanP. (1993). Improving generalization for temporal difference learning: the successor representation. Neural Comput.5, 613–624. doi: 10.1162/neco.1993.5.4.613
19
de CothiW.BarryC. (2020). Neurobiological successor features for spatial navigation. Hippocampus30, 1347–1355. doi: 10.1002/hipo.23246
20
De MartinoB.CorteseA. (2023). Goals, usefulness and abstraction in value-based choice. Trends Cogn. Sci.27, 65–80. doi: 10.1016/j.tics.2022.11.001
21
DeckerJ. H.OttoA. R.DawN. D.HartleyC. A. (2016). From creatures of habit to goal-directed learners: tracking the developmental emergence of model-based reinforcement learning. Psychol. Sci.27, 848–858. doi: 10.1177/0956797616639301
22
DoellerC. F.BarryC.BurgessN. (2010). Evidence for grid cells in a human memory network. Nature463, 657–661. doi: 10.1038/nature08704
23
DragoiG.TonegawaS. (2011). Preplay of future place cell sequences by hippocampal cellular assemblies. Nature469, 397–401. doi: 10.1038/nature09633
24
DragoiG.TonegawaS. (2013). Distinct preplay of multiple novel spatial experiences in the rat. Proc. Natl. Acad. Sci. U. S. A.110, 9100–9105. doi: 10.1073/pnas.1306031110
25
DubéL.SilveiraP. P.NielsenD. E.MooreS.PaquetC.Cisneros-FrancoJ. M.et al. (2022). From precision medicine to precision convergence for multilevel resilience—the aging brain and its social isolation. Front. Public Health10:720117. doi: 10.3389/fpubh.2022.720117
26
DuvelleÉ.GrievesR. M.van der MeerM. A. A. (2023). Temporal context and latent state inference in the hippocampal splitter signal. elife12, 1–35. doi: 10.7554/eLife.82357
27
EichenbaumH. (2017). On the integration of space, time, and memory. Neuron95, 1007–1018. doi: 10.1016/j.neuron.2017.06.036
28
EkstromA. D.KahanaM. J.CaplanJ. B.FieldsT. A.IshamE. A.NewmanE. L.et al. (2003). Cellular networks underlying human spatial navigation. Nature425, 184–188. doi: 10.1038/nature01964
29
EllisC. T.SkalabanL. J.YatesT. S.BejjankiV. R.CórdovaN. I.Turk-BrowneN. B. (2021). Evidence of hippocampal learning in human infants. Curr. Biol.31, 3358–3364.e4. doi: 10.1016/j.cub.2021.04.072
30
FangC.AronovD.AbbottL. F.MackeviciusE. (2023). Neural learning rules for generating flexible predictions and computing the successor representation. elife12:e80680. doi: 10.7554/eLife.80680
31
FarooqU.DragoiG. (2019). Emergence of preconfigured and plastic time-compressed sequences in early postnatal development. Science363, 168–173. doi: 10.1126/science.aav0502
32
FarooqU.SibilleJ.LiuK.DragoiG. (2019). Strengthened temporal coordination within pre-existing sequential cell assemblies supports trajectory replay. Neuron103, 719–733.e7. doi: 10.1016/j.neuron.2019.05.040
33
Feher da SilvaC.LombardiG.EdelsonM.HareT. A. (2023). Rethinking model-based and model-free influences on mental effort and striatal prediction errors. Nat. Hum. Behav.7, 956–969. doi: 10.1038/s41562-023-01573-1
34
FellowsL. K.FarahM. J. (2007). The role of ventromedial prefrontal cortex in decision making: judgment under uncertainty or judgment per se?Cereb. Cortex17, 2669–2674. doi: 10.1093/cercor/bhl176
35
FerbinteanuJ. (2020). The hippocampus and dorso-lateral striatum integrate distinct types of memories through time and space, respectively. J. Neurosci.40, 9055–9065. doi: 10.1523/JNEUROSCI.1084-20.2020
36
Fernández-RuizA.OlivaA.Fermino de OliveiraE.Rocha-AlmeidaF.TingleyD.BuzsákiG. (2019). Long-duration hippocampal sharp wave ripples improve memory. Science364, 1082–1086. doi: 10.1126/science.aax0758
37
FosterD. J.MorrisR. G. M.DayanP. (2000). A model of hippocampally dependent navigation, using the temporal difference learning rule. Hippocampus10, 1–16. doi: 10.1002/(SICI)1098-1063(2000)10:1<1::AID-HIPO1>3.0.CO;2-1
38
FosterD. J.WilsonM. A. (2007). Hippocampal theta sequences. Hippocampus17, 1093–1099. doi: 10.1002/hipo.20345
39
FrankL. M.BrownE. N.WilsonM. (2000). Trajectory encoding in the hippocampus and entorhinal cortex. Neuron27, 169–178. doi: 10.1016/S0896-6273(00)00018-0
40
GallagherS. (2000). Philosophical conceptions of the self: implications for cognitive science. Trends Cogn. Sci.4, 14–21. doi: 10.1016/S1364-6613(99)01417-5
41
GauthierJ. L.TankD. W. (2018). A dedicated population for reward coding in the hippocampus. Neuron99, 179–193.e7. doi: 10.1016/j.neuron.2018.06.008
42
GeertsJ. P.ChersiF.StachenfeldK. L.BurgessN. (2020). A general model of hippocampal and dorsal striatal learning and decision making. Proc. Natl. Acad. Sci. U. S. A.117, 31427–31437. doi: 10.1073/pnas.2007981117
43
GeorgeT. M.de CothiW.StachenfeldK.BarryC. (2023). Rapid learning of predictive maps with STDP and theta phase precession. elife12:e80663. doi: 10.7554/eLife.80663
44
GeorgeD.RikhyeR. V.GothoskarN.GuntupalliJ. S.DedieuA.Lázaro-GredillaM. (2021). Clone-structured graph representations enable flexible learning and vicarious evaluation of cognitive maps. Nat. Commun.12, 1–17. doi: 10.1038/s41467-021-22559-5
45
GershmanS. J. (2018). The successor representation: its computational logic and neural substrates. J. Neurosci.38, 7193–7200. doi: 10.1523/JNEUROSCI.0151-18.2018
46
GershmanS. J.HorvitzE. J.TenenbaumJ. B. (2015). Computational rationality: a converging paradigm for intelligence in brains, minds, and machines. Science349, 273–278. doi: 10.1126/science.aac6076
47
GershmanS. J.MooreC. D.ToddM. T.NormanK. A.SederbergP. B. (2012). The successor representation and temporal context. Neural Comput.24, 1553–1568. doi: 10.1162/NECO_a_00282
48
GirardeauG.BenchenaneK.WienerS. I.BuzsákiG.ZugaroM. B. (2009). Selective suppression of hippocampal ripples impairs spatial memory. Nat. Neurosci.12, 1222–1223. doi: 10.1038/nn.2384
49
GoodroeS. C.StarnesJ.BrownT. I. (2018). The complex nature of hippocampal-striatal interactions in spatial navigation. Front. Hum. Neurosci.12, 1–9. doi: 10.3389/fnhum.2018.00250
50
GuptaA. S.van der MeerM. A. A.TouretzkyD. S.RedishA. D. (2010). Hippocampal replay is not a simple function of experience. Neuron65, 695–705. doi: 10.1016/j.neuron.2010.01.034
51
HerbertC.BlumeC.NorthoffG. (2016). Can we distinguish an “I” and “ME” during listening?—an event-related EEG study on the processing of first and second person personal and possessive pronouns. Self Identity15, 120–138. doi: 10.1080/15298868.2015.1085893
52
HollupS. A.MoldenS.DonnettJ. G.MoserM. B.MoserE. I. (2001). Accumulation of hippocampal place fields at the goal location in an annular watermaze task. J. Neurosci.21, 1635–1644. doi: 10.1523/JNEUROSCI.21-05-01635.2001
53
HowardM. W.KahanaM. J. (2002). A distributed representation of temporal context. J. Math. Psychol.46, 269–299. doi: 10.1006/jmps.2001.1388
54
HowardM. W.MacDonaldC. J.TiganjZ.ShankarK. H.duQ.HasselmoM. E.et al. (2014). A unified mathematical framework for coding time, space, and sequences in the hippocampal region. J. Neurosci.34, 4692–4707. doi: 10.1523/JNEUROSCI.5808-12.2014
55
HowardM. W.ShankarK. H.AueW. R.CrissA. H. (2015). A distributed representation of internal time. Psychol. Rev.122, 24–53. doi: 10.1037/a0037840
56
IgataH.IkegayaY.SasakiT. (2021). Prioritized experience replays on a hippocampal predictive map for learning. Proc. Natl. Acad. Sci. U. S. A.118, 1–9. doi: 10.1073/pnas.2011266118
57
ItoH. T.ZhangS. J.WitterM. P.MoserE. I.MoserM. B. (2015). A prefrontal–thalamo–hippocampal circuit for goal-directed spatial navigation. Nature522, 50–55. doi: 10.1038/nature14396
58
ItoR.LeeA. C. H. (2016). The role of the hippocampus in approach-avoidance conflict decision-making: evidence from rodent and human studies. Behav. Brain Res.313, 345–357. doi: 10.1016/j.bbr.2016.07.039
59
JacobacciF.ArmonyJ. L.YeffalA.LernerG.AmaroE.JovicichJ.et al. (2020). Rapid hippocampal plasticity supports motor sequence learning. Proc. Natl. Acad. Sci. U. S. A.117, 23898–23903. doi: 10.1073/pnas.2009576117
60
JacobsJ.WeidemannC. T.MillerJ. F.SolwayA.BurkeJ. F.WeiX. X.et al. (2013). Direct recordings of grid-like neuronal activity in human spatial navigation. Nat. Neurosci.16, 1188–1190. doi: 10.1038/nn.3466
61
JankowskiM. M.IslamM. N.WrightN. F.VannS. D.ErichsenJ. T.AggletonJ. P.et al. (2014). Nucleus reuniens of the thalamus contains head direction cells. elife3, 1–10. doi: 10.7554/eLife.03075
62
JohnsonA.RedishA. D. (2007). Neural ensembles in CA3 transiently encode paths forward of the animal at a decision point. J. Neurosci.27, 12176–12189. doi: 10.1523/JNEUROSCI.3761-07.2007
63
JohnsonA.van der MeerM. A.RedishA. D. (2007). Integrating hippocampus and striatum in decision-making. Curr. Opin. Neurobiol.17, 692–697. doi: 10.1016/j.conb.2008.01.003
64
JohnsonA.VendittoS. (2015). “Reinforcement learning and hippocampal dynamics” in Analysis and modeling of coordinated multi-neuronal activity. Springer series in computational neuroscience. ed. TatsunoM., vol. 12 (New York, NY: Springer), 299–312.
65
JungM. W.WienerS. I.McNaughtonB. L. (1994). Comparison of spatial firing characteristics of units in dorsal and ventral hippocampus of the rat. J Neurosci.14, 7347–56. doi: 10.1523/JNEUROSCI.14-12-07347.1994
66
KaminL. J. (1969). “Predictability, surprise, attention, and conditioning,” in Punishment and Aversive Behavior. eds. CampbellB. A.ChurchR. M. (New York: Appleton-Century-Crofts), 279–296.
67
KaplanR.AdhikariM. H.HindriksR.MantiniD.MurayamaY.LogothetisN. K.et al. (2016). Hippocampal sharp-wave ripples influence selective activation of the default mode network. Curr. Biol.26, 686–691. doi: 10.1016/j.cub.2016.01.017
68
KayK.ChungJ. E.SosaM.SchorJ. S.KarlssonM. P.LarkinM. C.et al. (2020). Constant sub-second cycling between representations of possible futures in the hippocampus. Cells180, 552–567.e25. doi: 10.1016/j.cell.2020.01.014
69
KennerleyS. W.WaltonM. E. (2011). Decision making and reward in frontal cortex: complementary evidence from neurophysiological and neuropsychological studies. Behav. Neurosci.125, 297–317. doi: 10.1037/a0023575
70
KimJ.GhimJ. W.LeeJ. H.JungM. W. (2013). Neural correlates of interval timing in rodent prefrontal cortex. J. Neurosci.33, 13834–13847. doi: 10.1523/JNEUROSCI.1443-13.2013
71
KjelstrupK. B.SolstadT.BrunV. H.HaftingT.LeutgebS.WitterM. P.et al. (2008). Finite scale of spatial representation in the hippocampus. Science321, 140–143. doi: 10.1126/science.1157086
72
KnudsenE. B.WallisJ. D. (2021). Hippocampal neurons construct a map of an abstract value space. Cells184, 4640–4650.e10. doi: 10.1016/j.cell.2021.07.010
73
KobanL.GianarosP. J.KoberH.WagerT. D. (2021). The self in context: brain systems linking mental and physical health. Nat. Rev. Neurosci.22, 309–322. doi: 10.1038/s41583-021-00446-8
74
KotchoubeyB.TretterF.BraunH. A.BuchheimT.DraguhnA.FuchsT.et al. (2016). Methodological problems on the way to integrative human neuroscience. Front. Integr. Neurosci.10, 1–19. doi: 10.3389/fnint.2016.00041
75
KowadloG.AhmedA.MayanA.RawlinsonD. (2022). Continual few-shot learning with hippocampal-inspired replay. arxiv:2209.07863. doi: 10.48550/arXiv.2209.07863
76
KruglanskiA. W.SzumowskaE. (2020). Habitual behavior is goal-driven. Perspect. Psychol. Sci.15, 1256–1271. doi: 10.1177/1745691620917676
77
LeeS.YuL. Q.LermanC.KableJ. W. (2021). Subjective value, not a gridlike code, describes neural activity in ventromedial prefrontal cortex during value-based decision-making. Neuroimage237:118159. doi: 10.1016/j.neuroimage.2021.118159
78
LeonM. I.ShadlenM. N. (2003). Representation of time by neurons in the posterior parietal cortex of the macaque. Neuron38, 317–327. doi: 10.1016/S0896-6273(03)00185-5
79
LinL.-J. (1992). Self-improving reactive agents based on reinforcement learning, planning and teaching. Mach. Learn.8, 293–321. doi: 10.1007/BF00992699
80
LipsmanN.NakaoT.KanayamaN.KraussJ. K.AndersonA.GiacobbeP.et al. (2014). Neural overlap between resting state and self-relevant activity in human subcallosal cingulate cortex – single unit recording in an intracranial study. Cortex60, 139–144. doi: 10.1016/j.cortex.2014.09.008
81
LismanJ. E.GraceA. A. (2005). The hippocampal-VTA loop: controlling the entry of information into long-term memory. Neuron46, 703–713. doi: 10.1016/j.neuron.2005.05.002
82
LismanJ.RedishA. D. (2009). Prediction, sequences and the hippocampus. Philos. Trans. R. Soc. Lond.364, 1193–1201. doi: 10.1098/rstb.2008.0316
83
LiuY.DolanR. J.Kurth-NelsonZ.BehrensT. E. J. (2019). Human replay spontaneously reorganizes experience. Cells178, 640–652.e14. doi: 10.1016/j.cell.2019.06.012
84
LustigC.MatellM. S.MeckW. H. (2005). Not “just” a coincidence: frontal-striatal interactions in working memory and interval timing. Memory13, 441–448. doi: 10.1080/09658210344000404
85
MaingretN.GirardeauG.TodorovaR.GoutierreM.ZugaroM. (2016). Hippocampo-cortical coupling mediates memory consolidation during sleep. Nat. Neurosci.19, 959–964. doi: 10.1038/nn.4304
86
MannsJ. R.HowardM. W.EichenbaumH. (2007). Gradual changes in hippocampal activity support remembering the order of events. Neuron56, 530–540. doi: 10.1016/j.neuron.2007.08.017
87
MarrD. (1971). Simple memory: a theory for achicortex. Philos. Trans. R. Soc. B Biol. Sci.262, 23–81. doi: 10.1098/rstb.1971.0078
88
MattarM. G.LengyelM. (2022). Planning in the brain. Neuron110, 914–934. doi: 10.1016/j.neuron.2021.12.018
89
MillerK. J.BotvinickM. M.BrodyC. D. (2017). Dorsal hippocampus contributes to model-based planning. Nat. Neurosci.20, 1269–1276. doi: 10.1038/nn.4613
90
MillerK. J.BotvinickM. M.BrodyC. D. (2022). Value representations in the rodent orbitofrontal cortex drive learning, not choice. elife11, 1–27. doi: 10.7554/eLife.64575
91
MızrakE.BouffardN. R.LibbyL. A.BoormanE. D.RanganathC. (2021). The hippocampus and orbitofrontal cortex jointly represent task structure during memory-guided decision making. Cell Rep.37:110065. doi: 10.1016/j.celrep.2021.110065
92
MomennejadI.HowardM. W. (2018). Predicting the future with multi-scale successor representations. bioRxiv.:449470. doi: 10.1101/449470
93
MomennejadI.OttoA. R.DawN. D.NormanK. A. (2018). Offline replay supports planning in human reinforcement learning. elife7, 1–25. doi: 10.7554/eLife.32548
94
MomennejadI.RussekE. M.CheongJ. H.BotvinickM. M.DawN. D.GershmanS. J. (2017). The successor representation in human reinforcement learning. Nat. Hum. Behav.1, 680–692. doi: 10.1038/s41562-017-0180-8
95
MullerR. U.KubieJ. L. (1987). The effects of changes in the environment on the spatial firing of hippocampal complex-spike cells. J. Neurosci.7, 1951–1968. doi: 10.1523/JNEUROSCI.07-07-01951.1987
96
NiehE. H.SchottdorfM.FreemanN. W.LowR. J.LewallenS.KoayS. A.et al. (2021). Geometry of abstract learned knowledge in the hippocampus. Nature595, 80–84. doi: 10.1038/s41586-021-03652-7
97
NormanY.YeagleE. M.KhuvisS.HarelM.MehtaA. D.MalachR. (2019). Hippocampal sharp-wave ripples linked to visual episodic recollection in humans. Science365:eaax1030. doi: 10.1126/science.aax1030
98
NorthoffG.HayesD. J. (2011). Is our self nothing but reward?Biol. Psychiatry69, 1019–1025. doi: 10.1016/j.biopsych.2010.12.014
99
O’KeefeJ. (1976). Place units in the hippocampus of the freely moving rat. Exp. Neurol.51, 78–109. doi: 10.1016/0014-4886(76)90055-8
100
O’KeefeJ.BurgessN. (2005). Dual phase and rate coding in hippocampal place cells: theoretical significance and relationship to entorhinal grid cells. Hippocampus15, 853–866. doi: 10.1002/hipo.20115
101
O’KeefeJ.RecceM. L. (1993). Phase relationship between hippocampal place units and the EEG theta rhythm. Hippocampus3, 317–330. doi: 10.1002/hipo.450030307
102
O’NeilE. B.NewsomeR. N.LiI. H. N.ThavabalasingamS.ItoR.LeeA. C. H. (2015). Examining the role of the human hippocampus in approach–avoidance decision making using a novel conflict paradigm and multivariate functional magnetic resonance imaging. J. Neurosci.35, 15039–15049. doi: 10.1523/JNEUROSCI.1915-15.2015
103
ÓlafsdóttirH. F.BarryC.SaleemA. B.HassabisD.SpiersH. J. (2015). Hippocampal place cells construct reward related sequences through unexplored space. elife4, 1–17. doi: 10.7554/eLife.06063
104
PastalkovaE.ItskovV.AmarasinghamA.BuzsákiG. (2008). Internally generated cell assembly sequences in the rat hippocampus. Science321, 1322–1327. doi: 10.1126/science.1159775
105
PavlovI. P. (1960). Conditioned reflex: an investigation of the physiological activity of the cerebral cortex. Oxford, England: Dover Publications, xi, 430.
106
PetersenP. C.BuzsákiG. (2020). Cooling of medial septum reveals theta phase lag coordination of hippocampal cell assemblies. Neuron107, 731–744.e3. doi: 10.1016/j.neuron.2020.05.023
107
PfeifferB. E.FosterD. J. (2013). Hippocampal place-cell sequences depict future paths to remembered goals. Nature497, 74–79. doi: 10.1038/nature12112
108
PowerJ. D.FairD. A.SchlaggarB. L.PetersenS. E. (2010). The development of human functional brain networks. Neuron67, 735–748. doi: 10.1016/j.neuron.2010.08.017
109
QasimS.MillerJ.InmanC. S.GrossR. E.WillieJ. T.LegaB.et al. (2018). Neurons remap to represent memories in the human entorhinal cortex. bioRxiv.:433862. doi: 10.1101/433862
110
RahwanI.CebrianM.ObradovichN.BongardJ.BonnefonJ. F.BreazealC.et al. (2019). Machine behaviour. Nature568, 477–486. doi: 10.1038/s41586-019-1138-y
111
RedishA. D. (2016). Vicarious trial and error. Nat. Rev. Neurosci.17, 147–159. doi: 10.1038/nrn.2015.30
112
RossR. S.SherrillK. R.SternC. E. (2011). The hippocampus is functionally connected to the striatum and orbitofrontal cortex during context dependent decision making. Brain Res.1423, 53–66. doi: 10.1016/j.brainres.2011.09.038
113
RussekE. M.MomennejadI.BotvinickM. M.GershmanS. J.DawN. D. (2017). Predictive representations can link model-based reinforcement learning to model-free mechanisms. PLoS Comput. Biol.13, 1–35. doi: 10.1371/journal.pcbi.1005768
114
SamsonR. D.FrankM. J.FellousJ. M. (2010). Computational models of reinforcement learning: the role of dopamine as a reward signal. Cogn. Neurodyn.4, 91–105. doi: 10.1007/s11571-010-9109-x
115
SarelA.FinkelsteinA.LasL.UlanovskyN. (2017). Vectorial representation of spatial goals in the hippocampus of bats. Science355, 176–180. doi: 10.1126/science.aak9589
116
SchaeferM.NorthoffG. (2017). Who am I: the conscious and the unconscious self. Front. Hum. Neurosci.11, 1–5. doi: 10.3389/fnhum.2017.00126
117
SchapiroA. C.McDevittE. A.RogersT. T.MednickS. C.NormanK. A. (2018). Human hippocampal replay during rest prioritizes weakly learned information and predicts memory performance. Nat. Commun.9, 1–11. doi: 10.1038/s41467-018-06213-1
118
SchuckN. W.NivY. (2019). Sequential replay of non-spatial task states in the human hippocampus HHS public access. Science364, 1–24. doi: 10.1126/science.aaw5181
119
SchultzW.DayanP.MontagueP. R. (1997). A neural substrate of prediction and reward. Science275, 1593–1599. doi: 10.1126/science.275.5306.1593
120
ScovilleW. B.MilnerB. (1957). Loss of recent memory after bilateral hippocampal lesions. 1957. J. Neuropsychiatry Clin. Neurosci.20, 11–21. doi: 10.1136/jnnp.20.1.11
121
SharpP. A.LangerR. (2011). Promoting convergence in biomedical science. Science333:527. doi: 10.1126/science.1205008
122
SimonD. A.DawN. D. (2011). Neural correlates of forward planning in a spatial decision task in humans. J. Neurosci.31, 5526–5539. doi: 10.1523/JNEUROSCI.4647-10.2011
123
SjulsonL.PeyracheA.CumpelikA.CassataroD.BuzsákiG. (2018). Cocaine place conditioning strengthens location-specific hippocampal coupling to the nucleus accumbens. Neuron98, 926–934.e5. doi: 10.1016/j.neuron.2018.04.015
124
SkaggsW. E.McNaughtonB. L. (1998). Spatial firing properties of hippocampal CA1 populations in an environment containing two visually identical regions. J. Neurosci.18, 8455–8466. doi: 10.1523/JNEUROSCI.18-20-08455.1998
125
SkaggsW. E.McNaughtonB. L.WilsonM. A.BarnesC. A. (1996). Theta phase precession in hippocampal neuronal populations and the compression of temporal sequences. Hippocampus6, 149–172. doi: 10.1002/(SICI)1098-1063(1996)6:2<149::AID-HIPO6>3.0.CO;2-K
126
SpallaD.CornacchiaI. M.TrevesA. (2021). Continuous attractors for dynamic memories. elife10, 1–28. doi: 10.7554/eLife.69499
127
StachenfeldK. L.BotvinickM. M.GershmanS. J. (2017). The hippocampus as a predictive map. Nat. Neurosci.20, 1643–1653. doi: 10.1038/nn.4650
128
StoianovI.MaistoD.PezzuloG. (2022). The hippocampal formation as a hierarchical generative model supporting generative replay and continual learning. Prog. Neurobiol.217:102329. doi: 10.1016/j.pneurobio.2022.102329
129
StoutJ. J.HallockH. L.GeorgeA. E.AdirajuS. S.GriffinA. L. (2022). The ventral midline thalamus coordinates prefrontal–hippocampal neural synchrony during vicarious trial and error. Sci. Rep.12, 1–13. doi: 10.1038/s41598-022-14707-8
130
TamuraM.SpellmanT. J.RosenA. M.GogosJ. A.GordonJ. A. (2017). Hippocampal-prefrontal theta-gamma coupling during performance of a spatial working memory task. Nat. Commun.8:2182. doi: 10.1038/s41467-017-02108-9
131
TeylerT. J.DiScennaP. (1986). The hippocampal memory indexing theory. Behav. Neurosci.100, 147–154. doi: 10.1037/0735-7044.100.2.147
132
TolmanE. C. (1948). Cognitive maps in rats and men. Psychol. Rev.55, 189–208. doi: 10.1037/h0061626
133
UmbachG.KantakP.JacobsJ.KahanaM.PfeifferB. E.SperlingM.et al. (2020). Time cells in the human hippocampus and entorhinal cortex support episodic memory. Proc. Natl. Acad. Sci. U. S. A.117, 28463–28474. doi: 10.1073/pnas.2013250117
134
UriaB.IbarzB.BaninoA.ZambaldiV.KumaranD.HassabisD.et al. (2020). The spatial memory pipeline: a model of egocentric to allocentric understanding in mammalian brains. bioRxiv.:378141. doi: 10.1101/2020.11.11.378141
135
van de VenG. M.SiegelmannH. T.ToliasA. S. (2020). Brain-inspired replay for continual learning with artificial neural networks. Nat. Commun.11:4069. doi: 10.1038/s41467-020-17866-2
136
VazA. P.WittigJ. H.Jr.InatiS. K.ZaghloulK. A. (2023). Backbone spiking sequence as a basis for preplay, replay, and default states in human cortex. Nat. Commun.14, 1–12. doi: 10.1038/s41467-023-40440-5
137
VertesR. P.HooverW. B.Szigeti-BuckK.LeranthC. (2007). Nucleus reuniens of the midline thalamus: link between the medial prefrontal cortex and the hippocampus. Brain Res. Bull.71, 601–609. doi: 10.1016/j.brainresbull.2006.12.002
138
ViardA.DoellerC. F.HartleyT.BirdC. M.BurgessN. (2011). Anterior Hippocampus and goal-directed spatial decision making. J. Neurosci.31, 4613–4621. doi: 10.1523/JNEUROSCI.4640-10.2011
139
VikbladhO. M.MeagerM. R.KingJ.BlackmonK.DevinskyO.ShohamyD.et al. (2019). Hippocampal contributions to model-based planning and spatial memory. Neuron102, 683–693.e4. doi: 10.1016/j.neuron.2019.02.014
140
WaeltiP.DickinsonA.SchultzW. (2001). Dopamine responses comply with basic assumptions of formal learning theory. Nature412, 43–48. doi: 10.1038/35083500
141
WeilbächerR. A.GluthS. (2017). The interplay of hippocampus and ventromedial prefrontal cortex in memory-based decision making. Brain Sci.7:4. doi: 10.3390/brainsci7010004
142
WhittingtonJ. C. R.McCaffaryD.BakermansJ. J. W.BehrensT. E. J. (2022). How to build a cognitive map. Nat. Neurosci.25, 1257–1272. doi: 10.1038/s41593-022-01153-y
143
WhittingtonJ. C. R.MullerT. H.MarkS.ChenG.BarryC.BurgessN.et al. (2020). The Tolman-Eichenbaum machine: unifying space and relational memory through generalization in the hippocampal formation. Cells183, 1249–1263.e23. doi: 10.1016/j.cell.2020.10.024
144
WikenheiserA. M.RedishA. D. (2015). Hippocampal theta sequences reflect current goals. Nat. Neurosci.18, 289–294. doi: 10.1038/nn.3909
145
WikenheiserA. M.SchoenbaumG. (2016). Over the river, through the woods: cognitive maps in the hippocampus and orbitofrontal cortex. Nat. Rev. Neurosci.17, 513–523. doi: 10.1038/nrn.2016.56
146
WilsonM. A.McNaughtonB. L. (1994). Reactivation of hippocampal ensemble memories during sleep. Science265, 676–679. doi: 10.1126/science.8036517
147
WilsonR. C.TakahashiY. K.SchoenbaumG.NivY. (2014). Orbitofrontal cortex as a cognitive map of task space. Neuron81, 267–279. doi: 10.1016/j.neuron.2013.11.005
148
WoodE. R.DudchenkoP. A.RobitsekR. J.EichenbaumH. (2000). Hippocampal neurons encode information about different types of memory episodes occurring in the same location. Neuron27, 623–633. doi: 10.1016/S0896-6273(00)00071-4
149
WoodW.MazarA.NealD. T. (2021). Habits and goals in human behavior: separate but interacting systems. Perspect. Psychol. Sci.17, 590–605. doi: 10.1177/1745691621994226
150
ZhangJ.HuangZ.ChenY.ZhangJ.GhindaD.NikolovaY.et al. (2018). Breakdown in the temporal and spatial organization of spontaneous brain activity during general anesthesia. Hum. Brain Mapp.39, 2035–2046. doi: 10.1002/hbm.23984
151
ZielinskiM. C.TangW.JadhavS. P. (2017). The role of replay and theta sequences in mediating hippocampal-prefrontal interactions. Hippocampus30, 60–72. doi: 10.1002/hipo.22821
Summary
Keywords
hippocampus, reinforcement learning, successor representation, decision making, self
Citation
Mehrotra D and Dubé L (2023) Accounting for multiscale processing in adaptive real-world decision-making via the hippocampus. Front. Neurosci. 17:1200842. doi: 10.3389/fnins.2023.1200842
Received
05 April 2023
Accepted
25 August 2023
Published
05 September 2023
Volume
17 - 2023
Edited by
Jochen Ditterich, University of California, Davis, United States
Reviewed by
Joshua Gold, University of Pennsylvania, United States; Nikita Sidorenko, University of Zurich, Switzerland; Petia D. Koprinkova-Hristova, Institute of Information and Communication Technologies (BAS), Bulgaria
Updates

Check for updates
Copyright
© 2023 Mehrotra and Dubé.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Dhruv Mehrotra, dhruv.mehrotra@mail.mcgill.ca
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.