Abstract
Reinforcement learning (RL) models have been influential in understanding many aspects of basal ganglia function, from reward prediction to action selection. Time plays an important role in these models, but there is still no theoretical consensus about what kind of time representation is used by the basal ganglia. We review several theoretical accounts and their supporting evidence. We then discuss the relationship between RL models and the timing mechanisms that have been attributed to the basal ganglia. We hypothesize that a single computational system may underlie both RL and interval timing—the perception of duration in the range of seconds to hours. This hypothesis, which extends earlier models by incorporating a time-sensitive action selection mechanism, may have important implications for understanding disorders like Parkinson's disease in which both decision making and timing are impaired.
Introduction
Computational models of reinforcement learning (RL) have had a profound influence on the contemporary understanding of the basal ganglia (Joel et al., ; Cohen and Frank, ). The central claim of these models is that the basal ganglia are organized to support prediction, learning and optimization of long-term reward. While this claim is now widely accepted, RL models have had little to say about the extensive research implicating the basal ganglia in interval timing—the perception of duration in the range of seconds to hours (Buhusi and Meck, ; Jones and Jahanshahi, ; Merchant et al., ). However, this is not to say that time is ignored by these models—on the contrary, time representation has been a pivotal issue in RL theory, particularly with regard to the role of dopamine (Suri and Schultz, 1999; Daw et al., ; Ludvig et al., ; Nakahara and Kaveri, ; Rivest et al., 2010).
In this review, we attempt a provisional synthesis of research on RL and interval timing in the basal ganglia. We begin by briefly reviewing RL models of the basal ganglia, with a focus on how they represent time. We then summarize the key data linking the basal ganglia with interval timing, drawing connections between computational approaches to timing and their relationship to RL models. Our central thesis is that by incorporating a time-sensitive action selection mechanism into RL models, a single computational system can support both RL and interval timing. This unified view leads to a coherent interpretation of decision making and timing deficits in Parkinson's disease.
Reinforcement learning models of the basal ganglia
RL models characterize animals as agents that seek to maximize future reward (for reviews, see Maia, ; Niv, ; Ludvig et al., ). To do so, animals are assumed to generate a prediction of future reward and select actions according to a policy that maximizes that reward. More formally, suppose that at time t an agent occupies a state st (e.g., the agent's location or the surrounding stimuli) and receives a reward rt. The agent's goal is to predict the expected discounted future return, or value, of visiting a sequence of states starting in state st (Sutton and Barto, 1998): where γ is a parameter that discounts distal rewards relative to proximal rewards, and E denotes an average over possibly stochastic sequences of states and rewards.
Typically, a state st is described by a set of D features, {xt (1), …, xt (D)}, encoding sensory and cognitive aspects of an animal's current experience. Given this state representation, the value can be approximated by a weighted combination of the features: where is an estimate of the true value V. According to RL models of the basal ganglia, these features are represented by cortical inputs to the striatum, with the striatum itself encoding the estimated value (Maia, ; Niv, ; Ludvig et al., ). The strengths of these corticostriatal synapses are represented by a set of weights {wt (1), …, wt (D)}.
These weights can be learned through a simple algorithm known as temporal-difference (TD) learning, which adjusts the weights on each time step based on the difference between received and predicted reward: where α is a learning rate and δt is a prediction error defined as:
The eligibility trace et(d) is updated according to: where λ is a decay parameter that determines the plasticity window of recent stimuli. The TD algorithm is a computationally efficient method that is known to converge to the true value function [see Equation (1) above] with enough experience and adequate features (Sutton and Barto, 1998).
The importance of this algorithm to neuroscience lies in the fact that the firing of midbrain dopamine neurons conforms remarkably well to the theoretical prediction error (Houk et al., ; Montague et al., ; Schultz et al., 1997; though see Redgrave et al., 2008 for a critique). For example, dopamine neurons increase their firing upon the delivery of an unexpected reward and pause when an expected reward is omitted (Schultz et al., 1997). The role of prediction errors in learning is supported by the observation that plasticity at corticostriatal synapses is gated by dopamine (Reynolds and Wickens, 2002; Steinberg et al., 2013), as well as a large body of behavioral evidence (Rescorla and Wagner, 1972; Sutton and Barto, 1990; Ludvig et al., ).
A fundamental question facing RL models is the choice of feature representation. Early applications of TD learning to the dopamine system assumed what is known as the complete serial compound (CSC; Moore et al., ; Sutton and Barto, 1990; Montague et al., ; Schultz et al., 1997), which represents every time step following stimulus onset as a separate feature. Thus, the first feature has a value of 1 for the first time step and 0 for all other time steps, the second feature has a value of 1 for the second time step and 0 for all other time steps, and so on. This CSC representation assumes a perfect clock, whereby the brain always knows exactly how many time steps have elapsed since stimulus onset.
The CSC is effective at capturing several salient aspects of the dopamine response to cued reward. A number of authors (e.g., Daw et al., ; Ludvig et al., ), however, have pointed out aspects of the dopamine response that appear inconsistent with the CSC. For example, the CSC predicts a large, punctate negative prediction error when an expected reward is omitted; the actual decrease in dopamine response is relatively small and temporally extended (Schultz et al., 1997; Bayer et al., ). Another problem with the CSC is that it predicts a large negative prediction error at the usual reward delivery time when a reward is delivered early. Contrary to this prediction, Hollerman and Schultz () found that early reward evoked a large response immediately after the unexpected reward, but showed little change from baseline at the usual reward delivery time.
It is possible that these mismatches between theory and data reflect problems with a number of different theoretical assumptions. Indeed, several theoretical assumptions have been questioned by recent research (see Niv, ). We focus here on alternative time representations as one potential response to the findings mentioned above.
We will discuss two of these alternatives (see also Suri and Schultz, 1999; Nakahara and Kaveri, ; Rivest et al., 2010): (1) the microstimulus representation and (2) states with variable durations (a semi-Markov formalism) and only partial observability. For the former, Ludvig et al. () proposed that when a stimulus is presented, it leaves a slowly decaying memory trace, which is encoded by a series of temporal receptive fields. Each feature (or “microstimulus”) xt(d) represents the proximity between the trace and the center of the receptive field, producing a spectrum of features that vary with time, as illustrated in Figure 1A. Specifically, Ludvig et al. endowed each stimulus with microstimuli of the following form: where D is the number of microstimuli, σ2 controls the width of each receptive field, and yt is the stimulus trace strength, which was set to 1 at stimulus onset and decreased exponentially with a decay rate of 0.985 per time step. Both cues and rewards elicit their own set of microstimuli. This feature representation is plugged into the TD learning equations described above.
Figure 1
The microstimulus representation is a temporally smeared version of the CSC: whereas in the CSC each feature encodes a single time point, in the microstimulus representation each feature encodes a temporal range (see also Grossberg and Schmajuk,
Recent data from Adler et al. (
A different solution to the limitations of the CSC was suggested by Daw et al. (
It is instructive to compare how these two models account for the data on early reward presented by Hollerman and Schultz (
Thus far, we have discussed time representations in the service of RL and their implications for the timing of the dopamine response during conditioning. What do RL models have to say about interval timing per se? We will argue below that these are not really separate problems: interval timing tasks can be viewed fundamentally as RL tasks. Concomitantly, the role of dopamine and the basal ganglia in interval timing can be understood in terms of their computational contributions to RL. To elaborate this argument, we need to first review some of the relevant theory and data linking interval timing with the basal ganglia.
Time representation in the basal ganglia: data and theory
The role of the basal ganglia and dopamine in interval timing has been studied most extensively in the context of two procedures: the peak procedure (Catania, 1970; Roberts, 1981) and the bisection procedure (Church and Deluty, 1977). The peak procedure consists of two trial types: on fixed-interval trials, the subject is rewarded if a response is made after a fixed duration following cue presentation. On probe trials, the cue duration is extended, and no reward is delivered for responding. Figure 2A shows a typical response curve on probe trials: on average, the response rate peaks around the time of food presentation (20 or 40 s in the figure) is ordinarily available and then decreases. The peak time (a measure of the animal's interval estimate) is the time at which the response rate is maximal.
Figure 2

Effects of methamphetamine on timing in (A) the peak procedure and (B) the bisection procedure in rats. Each curve in (A) represents response rate as a function of time, where time 0 corresponds to the trial onset. The methamphetamine curve corresponds to sessions in which rats were injected with methamphetamine; the baseline curve corresponds to sessions in which rats did not receive an injection. Each curve in (B) represents the proportion of trials on which the rat chose the “long” option as a function of probe cue duration. The saline curve corresponds to sessions in which the rat received a saline injection. In both procedures, methamphetamine leads to overestimation of the elapsing interval, producing early responding in the peak procedure and more “long” responses in the bisection procedure. Figure replotted from Maricq et al. (
The other two curves in Figure 2A illustrate the standard finding that drugs (or genetic manipulations) that increase dopamine transmission, such as methamphetamine, shift the response curve leftward (Maricq et al.,
In the bisection procedure, subjects are trained to respond differentially to short and long duration cues. Unreinforced probe trials with cue durations between these two extremes are occasionally presented. On these trials, typically, a psychometric curve is produced with greater selection of the long option (i.e., the option reinforced following long duration cues) with longer probes and greater selection of the short option with shorter probes and a gradual shift between the two (see Figure 2B). The indifference point or point of subjective equality is typically close to the geometric mean of the two anchor durations (Church and Deluty, 1977). Similar to the peak procedure, in the bisection procedure, Figure 2B shows how dopamine agonists usually produce a leftward shift in the psychometric curve—i.e., more “short” responses, whereas dopamine antagonists produce the opposite pattern (Maricq et al.,
The most influential interpretation of these findings draws upon the class of pacemaker-accumulator models (Gibbon et al.,
This interpretation is generally consistent with the findings from studies of patients with Parkinson's disease (PD), who have chronically low striatal dopamine levels. When off medication, these patients tend to underestimate the length of temporal intervals in verbal estimation tasks; dopaminergic medication alleviates this underestimation (Pastor et al., 1992; Lange et al.,
Pacemaker-accumulator models have been criticized on a number of grounds, such as lack of parsimony, implausible neurophysiological assumptions, and incorrect behavioral predictions (Staddon and Higa, 1999, 2006; Matell and Meck,
In fact, the microstimulus model of Ludvig et al. (
Toward a unified model of reinforcement learning and timing
One suggestive piece of evidence for how RL models and interval timing can be integrated comes from the study of Fiorillo et al. (
Figure 3

(A) Firing rates of dopamine neurons in monkeys to cues and rewards as a function of cue-reward interval duration. Adapted from Fiorillo et al. (
Whereas the response to the cue can be explained in terms of temporal discounting, the response to the reward should not (according to the CSC representation) depend on the cue-reward interval. The perfect timing inherent in the CSC representation means that the reward can be equally well predicted at all time points. Thus, there should be no reward-prediction error, and no phasic dopamine response, at the time of reward regardless of the cue-reward interval. Alternatively, the dopamine response to reward can be understood as reflecting increasing uncertainty in the temporal prediction. Figure 3B shows how, using the microstimulus TD model as defined as in Ludvig et al. (
Interval timing procedures, such as the peak procedure, add an additional nuance to this problem by introducing instrumental contingencies. Animals must now not only predict the timing of reward, but also learn when to respond. To analyze this problem in terms of RL, we need to augment the framework introduced earlier to have actions. There are various ways to accomplish this (see Sutton and Barto, 1998). The Actor-Critic architecture (Houk et al.,
When combined with the microstimulus representation, the actor-critic architecture naturally gives rise to timing behavior: in the peak procedure, on average, responding will tend to increase toward the expected reward time and decrease thereafter (see Figure 2). Importantly, the late microstimuli are less temporally precise than the early microstimuli, in the sense that their responses are more dispersed over time. As a consequence, credit for late rewards is assigned to a larger number of microstimuli. Under the assumption that response rate is proportion to predicted value, this dispersion of credit causes the timing of actions to be more spread out around the time of reward as the length of the interval increases, one of the central empirical regularities in timing behavior (Gibbon,
The partially observable semi-Markov model of Daw et al. (
Figure 4

(A) The distribution of prediction errors in the semi-Markov TD model (Daw et al.,
The microstimulus actor-critic model can also explain the effects of dopamine manipulations and Parkinson's disease. The key additional assumption is that early microstimuli (but not later ones) are primarily represented by the striatum. Timing in the milliseconds to seconds range depends on D2 receptors in the dorsal striatum (Rammsayer, 1993; Coull et al.,
A similar line of reasoning can explain some of the timing deficits in Parkinson's disease. The nigrostriatal pathway (the main source of dopamine to the dorsal striatum) is compromised in Parkinson's disease, resulting in reduced striatal dopamine levels. Because D2 receptors have a higher affinity for dopamine, Parkinson's disease leads to the predominance of D2-mediated activity and hence reduced striatal output (Wiecki and Frank, 2010). Our model thus predicts a rightward shift of estimated time, as is often observed experimentally (see above).
The linking of early microstimuli with the striatum in the model also leads to the prediction that low striatal dopamine levels will result in poorer learning of fast responses (which depend on the early microstimuli). In addition, responding will in general be slowed because the learned weights to the early microstimuli will be weak relative to those of late microstimuli. As a result, our model clearly predicts poorer learning of fast responses in Parkinson's disease. A study of temporal decision making in Parkinson's patients fits with this prediction (Moustafa et al., 2008). Patients were trained to respond at different latencies to a set of cues, with slow responses yielding more reward in an “increasing expected value” (IEV) condition and fast responses yielding more reward in a “decreasing expected value” (DEV) condition. It was found that the performance of medicated patients was better in the DEV condition, while performance of non-medicated patients was better in the IEV condition. If non-medicated patients have a paucity of short-timescale microstimuli (due to low striatal dopamine levels), then the model correctly anticipates that these patients will be impaired at learning about early events relative to later events.
Recently, Foerde et al. (
Conclusion
Timing and RL have for the most part been studied separately, giving rise to largely non-overlapping computational models. We have argued here, however, that these models do in fact share some important commonalities and reconciling them may provide a unified explanation of many behavioral and neural phenomena. While in this brief review we have only sketched such a synthesis, our goal is to plant the seeds for future theoretical unification.
One open question concerns how to reconcile the disparate theoretical ideas about time representation that were described in this paper. Our synthesis proposed a central role for a distributed elements representation of time such as the microstimuli of Ludvig et al. (
Conflict of interest statement
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Statements
Acknowledgments
We thank Marc Howard and Nathaniel Daw for helpful discussions. Samuel J. Gershman was supported by IARPA via DOI contract D10PC2002 and by a postdoctoral fellowship from the MIT Intelligence Initiative. Ahmed A. Moustafa is partially supported by a 2013 internal UWS Research Grant Scheme award P00021210. Elliot A. Ludvig was partially supported by NIH Grant #P30 AG024361 and the Princeton Pyne Fund.
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
References
1
AdlerA.KatabiS.FinkesI.IsraelZ.PrutY.BergmanH. (2012). Temporal convergence of dynamic cell assemblies in the striato-pallidal network. J. Neurosci. 32, 2473–2484. 10.1523/JNEUROSCI.4830-11.2012
2
ArtiedaJ.PastorM. A.LacruzF.ObesoJ. A. (1992). Temporal discrimination is abnormal in Parkinson's disease. Brain115, 199–210. 10.1093/brain/115.1.199
3
BalciF.LudvigE. A.AbnerR.ZhuangX.PoonP.BrunnerD. (2010). Motivational effects on interval timing in dopamine transporter (DAT) knockdown mice. Brain Res. 1325, 89–99. 10.1016/j.brainres.2010.02.034
4
BalciF.LudvigE. A.GibsonJ. M.AllenB. D.FrankK. M.KapustinskiB. J.et al. (2008). Pharmacological manipulations of interval timing using the peak procedure in male C3H mice. Psychopharmacology (Berl.)201, 67–80. 10.1007/s00213-008-1248-y
5
BayerH. M.GlimcherP. W. (2005). Midbrain dopamine neurons encode a quantitative reward prediction error signal. Neuron47, 129–141. 10.1016/j.neuron.2005.05.020
6
BayerH. M.LauB.GlimcherP. W. (2007). Statistics of midbrain dopamine neuron spike trains in the awake primate. J. Neurophysiol. 98, 1428–1439. 10.1152/jn.01140.2006
7
BornsteinA. M.DawN. D. (2012). Dissociating hippocampal and striatal contributions to sequential prediction learning. Eur. J. Neurosci. 35, 1011–1023. 10.1111/j.1460-9568.2011.07920.x
8
BuhusiC. V.MeckW. H. (2005). What makes us tick? Functional and neural mechanisms of interval timing. Nat. Rev. Neurosci. 6, 755–765. 10.1038/nrn1764
9
BuonomanoD. V.LajeR. (2010). Population clocks: motor timing with neural dyanmics. Trends Cogn. Sci. 14, 520–527. 10.1016/j.tics.2010.09.002
10
CataniaR. A. (1970). Reinforcement schedules and psychophysical judgements, in The Theory of Reinforcement Schedules, ed SchoenfeldW. N. (New York, NY: Appleton-Century Crofts), 1–42.
11
ChengR. K.AliY. M.MeckW. H. (2007). Ketamine “unlocks” the reduced clock-speed effects of cocaine following extended training: evidence for dopamine-glutamate interactions in timing and time perception. Neurobiol. Learn. Mem. 88, 149–159. 10.1016/j.nlm.2007.04.005
12
ChurchR. M.DelutyH. Z. (1977). The bisection of temporal intervals. J. Exp. Psychol. Anim. Behav. Process. 3, 216–228. 10.1037/0097-7403.3.3.216
13
CohenM. X.FrankM. J. (2009). Neurocomputational models of basal ganglia function in learning, memory and choice. Behav. Brain Res. 199, 141–156. 10.1016/j.bbr.2008.09.029
14
CoullJ. T.ChengR. K.MeckW. H. (2011). Neuroanatomical and neurochemical substrates of timing. Neuropsychopharmacology36, 3–25. 10.1038/npp.2010.113
15
DawN. D.CourvilleA. C.TouretzkyD. S. (2006). Representation and timing in theories of the dopamine system. Neural Comput. 18, 1637–1677. 10.1162/neco.2006.18.7.1637
16
DawN. D.KakadeS.DayanP. (2002). Opponent interactions between serotonin and dopamine. Neural Netw. 15, 603–616. 10.1016/S0893-6080(02)00052-7
17
DrewM. R.FairhurstS.MalapaniC.HorvitzJ. C.BalsamP. D. (2003). Effects of dopamine antagonists on the timing of two intervals. Pharmacol. Biochem. Behav. 75, 9–15. 10.1016/S0091-3057(03)00036-4
18
FiorilloC. D.NewsomeW. T.SchultzW. (2008). The temporal precision of reward prediction in dopamine neurons. Nat. Neurosci. 11, 966–973. 10.1038/nn.2159
19
FoerdeK.RaceE.VerfaellieM.ShohamyD. (2013). A role for the medial temporal lobe in feedback-driven learning: evidence from amnesia. J. Neurosci. 33, 5698–5704. 10.1523/JNEUROSCI.5217-12.2013
20
GerfenC. R. (1992). The neostriatal mosaic: multiple levels of compartmental organization in the basal ganglia. Annu. Rev. Neurosci. 15, 285–320. 10.1146/annurev.ne.15.030192.001441
21
GibbonJ. (1977). Scalar expectancy theory and Weber's law in animal timing. Psychol. Rev. 84, 279–325. 10.1037/0033-295X.84.3.279
22
GibbonJ.ChurchR. M.MeckW. H. (1984). Scalar timing in memory, in Annals of the New York Academy of Sciences: Timing and Time Perception, Vol. 423, eds GibbonJ.AllanL. G. (New York, NY: New York Academy of Sciences), 52–77.
23
GibbonJ.MalapaniC.DaleC. L.GallistelC. R. (1997). Towards a neurobiology of temporal cognition: advances and challenges. Curr. Opin. Neurobiol. 7, 170–184. 10.1016/S0959-4388(97)80005-0
24
GrossbergS.SchmajukN. A. (1989). Neural dynamics of adaptive timing and temporal discrimination during associative learning. Neural Netw. 2, 79–102. 10.1016/0893-6080(89)90026-9
25
HollermanJ. R.SchultzW. (1998). Dopamine neurons report an error in the temporal prediction of reward during learning. Nat. Neurosci. 1, 304–309. 10.1038/1124
26
HoukJ. C.AdamsJ. L.BartoA. G. (1995). A model of how the basal ganglia generate and use neural signals that predict reinforcement, in Models of information Processing in the Basal Ganglia, eds HoukJ. C.DavisJ. L.BeiserD. G. (Cambridge, MA: MIT Press), 249–270.
27
JinD. Z.FujiiN.GraybielA. M. (2009). Neural representation of time in cortico-basal ganglia circuits. Proc. Natl. Acad. Sci. U.S.A. 106, 19156–19161. 10.1073/pnas.0909881106
28
JoelD.NivY.RuppinE. (2002). Actor-critic models of the basal ganglia: new anatomical and computational perspectives. Neural Netw. 15, 535–547. 10.1016/S0893-6080(02)00047-3
29
JonesC. R.JahanshahiM. (2009). The substantia nigra, the basal ganglia, dopamine and temporal processing. J. Neural Transm. Suppl. 73, 161–171. 10.1007/978-3-211-92660-4_13
30
KobayashiS.SchultzW. (2008). Influence of reward delays on responses of dopamine neurons. J. Neurosci. 28, 7837–7846. 10.1523/JNEUROSCI.1600-08.2008
31
LangeK. W.TuchaO.SteupA.GsellW.NaumannM. (1995). Subjective time estimation in Parkinson's disease. J. Neural Transm. Suppl. 46, 433–438.
32
LeonM. I.ShadlenM. N. (2003). Representation of time by neurons in the posterior parietal cortex of the macaque. Neuron38, 317–327. 10.1016/S0896-6273(03)00185-5
33
LudvigE. A.BellemareM. G.PearsonK. G. (2011). A primer on reinforcement learning in the brain: psychological, computational, and neural perspectives, in Computational Neuroscience for Advancing Artificial Intelligence: Models, Methods and Applications, eds AlonsoE.MondragonE. (Hershey, PA: IGI Global), 111–144.
34
LudvigE. A.SuttonR. S.KehoeE. J. (2008). Stimulus representation and the timing of reward-prediction errors in models of the dopamine system. Neural Comput. 20, 3034–3054. 10.1162/neco.2008.11-07-654
35
LudvigE. A.SuttonR. S.KehoeE. J. (2012). Evaluating the TD model of classical conditioning. Learn. Behav. 40, 305–319. 10.3758/s13420-012-0082-6
36
LudvigE. A.SuttonR. S.VerbeekE. L.KehoeE. J. (2009). A computational model of hippocampal function in trace conditioning. Adv. Neural Inf. Process. Syst. 21, 993–1000.
37
MacdonaldC. J.MeckW. H. (2005). Differential effects of clozapine and haloperidol on interval timing in the supraseconds range. Psychopharmacology (Berl.)182, 232–244. 10.1007/s00213-005-0074-8
38
MachadoA. (1997). Learning the temporal dynamics of behavior. Psychol. Rev. 104, 241–265. 10.1037/0033-295X.104.2.241
39
MaiaT. (2009). Reinforcement learning, conditioning, and the brain: successes and challenges. Cogn. Affect. Behav. Neurosci. 9, 343–364. 10.3758/CABN.9.4.343
40
MalapaniC.RakitinB.LevyR.MeckW. H.DeweerB.DuboisB.et al. (1998). Coupled temporal memories in Parkinson's disease: a dopamine-related dysfunction. J. Cogn. Neurosci. 10, 316–331. 10.1162/089892998562762
41
MaricqA. V.ChurchR. M. (1983). The differential effects of haloperidol and methamphetamine on time estimation in the rat. Psychopharmacology (Berl.)79, 10–15. 10.1007/BF00433008
42
MaricqA. V.RobertsS.ChurchR. M. (1981). Methamphetamine and time estimation. J. Exp. Psychol. Anim. Behav. Process. 7, 18–30. 10.1037//0097-7403.7.1.18
43
MatellM. S.BatesonM.MeckW. H. (2006). Single-trials analyses demonstrate that increases in clock speed contribute to the methamphetamine-induced horizontal shifts in peak-interval timing functions. Psychopharmacology (Berl.)188, 201–212. 10.1007/s00213-006-0489-x
44
MatellM. S.KingG. R.MeckW. H. (2004). Differential modulation of clock speed by the administration of intermittent versus continuous cocaine. Behav. Neurosci. 118, 150–156. 10.1037/0735-7044.118.1.150
45
MatellM. S.MeckW. H. (2004). Cortico-striatal circuits and interval timing: coincidence detection of oscillatory processes. Cogn. Brain Res. 21, 139–170. 10.1016/j.cogbrainres.2004.06.012
46
McClureE. A.SaulsgiverK. A.WynneC. D. L. (2005). Effects of d-amphetamine on temporal discrimination in pigeons. Behav. Pharmacol. 16, 193–208. 10.1097/01.fbp.0000171773.69292.bd
47
MeckW. H. (1986). Affinity for the dopamine D2 receptor predicts neuroleptic potency in decreasing the speed of an internal clock. Pharmacol. Biochem. Behav. 25, 1185–1189. 10.1016/0091-3057(86)90109-7
48
MerchantH.HarringtonD. L.MeckW. H. (2013). Neural basis of the perception and estimation of time. Annu. Rev. Neurosci. 36, 313–336. 10.1146/annurev-neuro-062012-170349
49
MiallC. (1989). The storage of time intervals using oscillating neurons. Neural Comput. 1, 359–371. 10.1162/neco.1989.1.3.359
50
MontagueP. R.DayanP.SejnowskiT. J. (1996). A framework for mesencephalic dopamine systems based on predictive Hebbian learning. J. Neurosci. 16, 1936–1947.
51
MooreJ. W.DesmondJ. E.BerthierN. E. (1989). Adaptively timed conditioned responses and the cerebellum: a neural network approach. Biol. Cybern. 62, 17–28. 10.1007/BF00217657
52
MoustafaA. A.CohenM. X.ShermanS. J.FrankM. J. (2008). A role for dopamine in temporal decision making and reward maximization in Parkinsonism. J. Neurosci. 28, 12294–12304. 10.1523/JNEUROSCI.3116-08.2008
53
NakaharaH.KaveriS. (2010). Internal-time temporal difference model for neural value-based decision making. Neural Comput. 22, 3062–3106. 10.1162/NECO_a_00049
54
NivY. (2009). Reinforcement learning in the brain. J. Math. Psychol. 53, 139–154. 10.1016/j.jmp.2008.12.005
55
OdumA. L.LievingL. M.SchaalD. W. (2002). Effects of D-amphetamine in a temporal discrimination procedure: selective changes in timing or rate dependency?J. Exp. Anal. Behav. 78, 195–214. 10.1901/jeab.2002.78-195
56
PastorM. A.ArtiedaJ.JahanshahiM.ObesoJ. A. (1992). Time estimation and reproduction is abnormal in Parkinson's disease. Brain115, 211–225. 10.1093/brain/115.1.211
57
RammsayerT. H. (1993). On dopaminergic modulation of temporal information processing. Biol. Psychol. 36, 209–222. 10.1016/0301-0511(93)90018-4
58
RedgraveP.GurneyK.ReynoldsJ. (2008). What is reinforced by phasic dopamine signals?Brain Res. Rev. 58, 322–339. 10.1016/j.brainresrev.2007.10.007
59
RescorlaR. A.WagnerA. R. (1972). A theory of Pavlovian conditioning: variations in the effectiveness of reinforcement and nonreinforcement, in Classical Conditioning II: Current Research and Theory, eds BlackA. H.ProkasyW. F. (New York, NY: Appleton-Century Crofts), 64–69.
60
ReynoldsJ. N. J.WickensJ. R. (2002). Dopamine-dependent plasticity of corticostriatal synapses. Neural Netw. 15, 507–521. 10.1016/S0893-6080(02)00045-X
61
RivestF.KalaskaJ. F.BengioY. (2010). Alternative time representation in dopamine models. J. Comput. Neurosci. 28, 107–130. 10.1007/s10827-009-0191-1
62
RobertsS. (1981). Isolation of an internal clock. J. Exp. Psychol. Anim. Behav. Process. 7, 242–268. 10.1037/0097-7403.7.3.242
63
SchultzW.DayanP.MontagueP. R. (1997). A neural substrate of prediction and reward. Science275, 1593–1599. 10.1126/science.275.5306.1593
64
ShankarK. H.HowardM. W. (2012). A scale-invariant internal representation of time. Neural Comput. 24, 134–193. 10.1162/NECO_a_00212
65
SimenP.BalciF.deSouzaL.CohenJ. D.HolmesP. (2011). A model of interval timing by neural integration. J. Neurosci. 31, 9238–9253. 10.1523/JNEUROSCI.3121-10.2011
66
SimenP.RivestF.LudvigE. A.BalciF.KilleenP. (2013). Timescale invariance in the pacemaker-accumulator family of timing models. Timing Time Percept. 1, 159–188. 10.1163/22134468-00002018
67
SpencerR. M.IvryR. B. (2005). Comparison of patients with Parkinson's disease or cerebellar lesions in the production of periodic movements involving event-based or emergent timing. Brain Cogn. 58, 84–93. 10.1016/j.bandc.2004.09.010
68
StaddonJ. E. R.HigaJ. J. (1999). Time and memory: towards a pacemaker-free theory of interval timing. J. Exp. Anal. Behav. 71, 215–251. 10.1901/jeab.1999.71-215
69
StaddonJ. E. R.HigaJ. J. (2006). Interval timing. Nat. Rev. Neurosci. 7. 10.1038/nrn1764-c1
70
SteinbergE. E.KeiflinR.BolvinJ. R.WittenI. B.DeisserothK.JanakP. H. (2013). A causal link between prediction errors, dopamine neurons and learning. Nat. Neurosci. 16, 966–973. 10.1038/nn.3413
71
SuriR. E.SchultzW. (1999). A neural network model with dopamine-like reinforcement signal that learns a spatial delayed response task. Neuroscience91, 871–890. 10.1016/S0306-4522(98)00697-6
72
SuttonR. S.BartoA. G. (1990). Time-derivative models of Pavlovian reinforcement, in Learning and Computational Neuroscience: Foundations of Adaptive Networks, eds GabrielM.MooreJ. (Cambridge, MA: MIT Press), 497–537.
73
SuttonR. S.BartoA. G. (1998). Reinforcement Learning: An Introduction. Cambridge, MA: MIT Press.
74
WeardenJ. H.Smith-SparkJ. H.CousinsR.EdelstynN. M.CodyF. W.O'BoyleD. J. (2008). Stimulus timing by people with Parkinson's disease. Brain Cogn. 67, 264–279. 10.1016/j.bandc.2008.01.010
75
WieckiT. V.FrankM. J. (2010). Neurocomputational models of motor and cognitive deficits in Parkinson's disease. Prog. Brain Res. 183, 275–297. 10.1016/S0079-6123(10)83014-6
Summary
Keywords
reinforcement learning, basal ganglia, dopamine, interval timing, Parkinson's disease
Citation
Gershman SJ, Moustafa AA and Ludvig EA (2014) Time representation in reinforcement learning models of the basal ganglia. Front. Comput. Neurosci. 7:194. doi: 10.3389/fncom.2013.00194
Received
15 October 2013
Accepted
23 December 2013
Published
09 January 2014
Volume
7 - 2013
Edited by
Hagai Bergman, The Hebrew University- Hadassah Medical School, Israel
Reviewed by
Yoram Burak, Hebrew University, Israel; Daoyun Ji, Baylor College of Medicine, USA
Copyright
© 2014 Gershman, Moustafa and Ludvig.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Samuel J. Gershman, Department of Brain and Cognitive Sciences, Massachusetts Institute of Technology, Room 46-4053, 77 Massachusetts Ave., Cambridge, MA 02139, USA e-mail: sjgershm@mit.edu
This article was submitted to the journal Frontiers in Computational Neuroscience.
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.