Abstract
Precommitment, or taking away a future choice from oneself, is a mechanism for overcoming impulsivity. Here we review recent work suggesting that precommitment can be best explained through a distributed decision-making system with multiple discounting rates. This model makes specific predictions about precommitment behavior and is especially interesting in light of the emerging multiple-systems view of decision-making, in which functional systems with distinct neural substrates use different computational strategies to optimize decisions. Given the growing consensus that impulsivity constitutes a common point of breakdown in decision-making processes, with common neural and computational mechanisms across multiple psychiatric disorders, it is useful to translate precommitment into the common language of temporal difference reinforcement learning that unites many of these behavioral and neural data.
It seems illogical on the surface, but humans and other animals sometimes put themselves in situations to prevent themselves from being given an option that they would choose if given the chance. They will even expend effort and cost to avoid being given the future option. Such restriction of one’s own future choices is called precommitment. It is theorized that precommitment occurs because humans and other animals have different preferences at different times (Strotz, 1955; Ainslie, ). Precommitment behaviors take many forms, ranging from purely external mechanisms like flushing cigarettes down the toilet, to purely internal mechanisms like making a promise to oneself that one is unwilling to break, to intermediate mechanisms like making a public statement about one’s intentions.
Precommitment is ubiquitous in human behavior. “Christmas Clubs,” popularized during the Great Depression, enforced saving through the year for Christmas shopping (Strotz, 1955). In the modern era, websites like stickk.com automatically transfer money from a credit card to a designated recipient (such as a charity) if the user fails to meet a specified goal (as reported by a trusted third party). In Australia, Canada, and Norway, many gambling machines require the gambler to pre-set a limit on his or her expenditure, after which the machine deactivates (Ladouceur et al., 2012). (Some gamblers also spontaneously create their own precommitment strategies, Wohl et al., 2008; Ladouceur et al., 2012.) In day-to-day experience, people place the ice cream out of sight, put money into a retirement account with withdrawal penalties, walk a different route to avoid seeing a store where there is temptation to buy something, or self-impose deadlines with self-imposed punishments (Ariely and Wertenbroch, ).
Precommitment behavior has been demonstrated in animals (Rachlin and Green, 1972; Ainslie, ), but there is not yet an established laboratory paradigm for eliciting precommitment behavior in humans. Although precommitment can be predicted to occur as a direct consequence of time-dependent changes in preference order (Ainslie, ), explicit neural and computational models of precommitment remain limited. In our paper, “A reinforcement learning model of precommitment in decision-making” (Kurth-Nelson and Redish, ), we examined whether current computational models of decision-making can explain precommitment and what those models imply for the mechanisms that underlie precommitment. Here, we will focus on integrating those results into the broader picture of decision-making.
Valuation and Discounting
Psychologists and economists (and now, neuroeconomists) operationalize the decision-making process through the framework of valuation. Whenever an organism (which we will call an “agent” here, to allow for easy translation between simulations and real organisms) is faced with a choice, each possible outcome is assigned a value. These values are compared, and the outcomes with higher values are more likely to be chosen (Glimcher, ). Although there are additional action-selection systems which do not work this way (such as reflexes), there is a compelling body of evidence that valuation plays a role in the making of many choices. Neural correlates of value-based decision-making have been identified in many parts of the brain (Rangel et al., 2008; Kable and Glimcher, ).
Rewards become less valued as they are more delayed – a phenomenon known as temporal or delay discounting. A discounting function is a quantitative description of this decay in value (Ainslie, ; Mazur, 1997; Madden and Bickel, 2010). The discounting function of an individual human subject can be measured empirically with a series of questions (for example, “Would you prefer $30 today or $100 in a year?”), and is generally stable over time (Ohmura et al., 2006; Takahashi et al., 2007; Jimura et al., ).
The simplest discounting function is one that decays exponentially. In exponential discounting, each unit of delay reduces value by the same percentage. However, when measured empirically, the discounting functions of humans and animals are not exponential (Ainslie, ; Madden and Bickel, 2010). Instead, they are steeper than exponential at short delays, and shallower than exponential at long delays (Figure 1). Hyperbolic functions are often used to fit these curves, but for our purposes it is not critical whether the shape is actually hyperbolic; only that it is more concave than exponential. All non-exponential functions show preference reversals (Strotz, 1955; Frederick et al., ) – an option preferred today is not necessarily preferred tomorrow.
Figure 1
Precommitment can be explained as a consequence of preference reversal (Ainslie,
Computational Models of Precommitment
Temporal difference reinforcement learning (TDRL) is often used to bridge the gap between descriptive theoretical models of decision-making and their neural implementation. Because of its biological plausibility, guaranteed convergence, and power to explain behavior and neural activity (Schultz et al., 1997; Sutton and Barto, 1998; Roesch et al., 2012), TDRL has become a well-established model of value-based decision-making (Montague et al., 1996; Schultz, 1998).
TDRL assumes that an agent can take actions, some of which are rewarded. The goal is to learn to take actions that maximize the reward received (Sutton and Barto, 1998). Distinct situations of the world are represented as states. TDRL aims to estimate the value of each state, which is defined as the total discounted future reward expected from that state. This is a recursive definition: the value of a state can be defined as the discounted value of the next state plus the reward available in the next state (Bellman,
In the standard implementation of TDRL, there is a state transition on every time step. Exponential discounting can therefore be calculated very straightforwardly by taking the value of the current state to be the value of the next state (plus the reward received if any) times a constant γ (0 < γ < 1). In this formulation, each unit of time causes the same attenuation of value, which is the definition of exponential discounting. However, non-exponential discounting has been difficult to implement in TDRL. There have been a handful of attempts at performing non-exponential (specifically, hyperbolic) discounting within a TDRL model (Daw,
We found that three of these four models produced hyperbolic discounting only in special cases (either across a single state transition, or in an environment with no choices) and therefore were unable to produce precommitment. The other model produced hyperbolic discounting in arbitrary state-spaces and was able to produce precommitment. The successfully precommitting model was the μAgents model that we introduced in 2009 – in this model, a set of exponentially discounting TDRL agents operating in parallel, each with a different discounting rate, and each maintaining its own estimate of the value function, collectively approximate hyperbolic discounting behavior (Figure 2; Kurth-Nelson and Redish,
Figure 2

Distributed discounting enables precommitment in temporal difference learning. (A) Twenty exponential curves with discounting rates spread uniformly between 0 and 1 are shown in black. The average of these curves is shown in red. This average curve closely approximates a hyperbolic function. (B) Standard TD models cannot precommit because, at each state transition, discounting starts over, ensuring that if SS is preferred over LL at the time of C, then it is also preferred at the time of P (top pair of curves). When averaging a set of exponential discount curves, discounting is not reset at each state transition, so preferences can reverse between C and P (bottom pair of curves).
A TDRL model of precommitment gives us a concrete computational hypothesis with which to explore potential mechanisms by which people choose to precommit. More generally, it is also important to have computational models that describe choice in complex state-spaces (Kurth-Nelson and Redish,
Predictions about Precommitment Behavior
Computational models allow exploration of parameter spaces. Although the fact that non-exponential (e.g., hyperbolic) discounting leads to preference reversals (Strotz, 1955; Frederick et al.,
First, the theoretical model predicts that precommitment is increased when there is a larger contrast between the SS and LL options. In other words, precommitment will be more favored if LL is very large and very delayed, compared to SS (of course, if LL is very large but not very delayed, then it will simply be preferred over SS at any time point, and precommitment will not be required). This suggests that, in the case of addiction, if we want to encourage precommitment, it is important to define the perceived alternative to drug use as being a major outcome, such as the long-term health and safety of oneself or family members (Heyman,
Second, we can predict that there is a complex effect of an agent’s discounting rate on their ability to precommit. When an agent is highly impulsive (fast discounting rate), it will be highly sensitive to the delay between precommitment and choice. If this delay is small, precommitment is unfavorable, but as this delay increases, the preference for precommitment increases steeply. On the other hand, if an agent is relatively patient (slow-discounting rate), then it will be largely insensitive to the delay between precommitment and choice, exhibiting at best a mild preference for precommitment for any value of this delay. Thus, the highest overall preference for precommitment appears in the most impulsive agents. On the surface this appears a bit paradoxical: the people with the strongest preference for an impulsive choice are the ones most likely to employ a strategy that curtails their ability to reach it. However, this finding suggests that in addiction, treatment strategies should be tailored to the individual depending on his or her own discounting rate. For fast discounters, inserting more time between precommitment and choice is essential – while for slow discounters, the theory predicts that it won’t make much of a difference. In fact, for slow-discounting addicts, precommitment may not be a useful strategy at all.
Third, the model predicts that precommitment is highly sensitive to the precise shape of an agent’s discounting function (Figure 3). Our theoretical analysis reveals that two discounting functions that are both fit by nearly identical hyperbolic parameters can exhibit entirely different patterns of precommitment behavior. In particular, the simulations in Kurth-Nelson and Redish (
Figure 3

Shape of discounting curve strongly influences precommitment. (A) The actual discounting curves of two individuals are shown in solid lines, and the best-fit hyperbolic curves are shown in dashed lines. These two subjects were both fit by a hyperbolic function with ln(K) of approximately 0 (from a range of −13 to +4 across subjects). (B) Predicted precommitment behavior, based on actual discounting curve shape of each subject, using the following parameters: DC = 6 days, DL = 1 day, DS = 0, RL = $150, RS = $100. Subject 1 is expected to have a modest preference for SS over LL, and to be averse to precommitment. Meanwhile, subject 2 is expected to have a strong preference for SS over LL, but to favor precommitment. (Data from Chopra et al.,
Multiple Systems
As noted above, TDRL models are incomplete descriptions of the full range of animal (including human) behavior (O’Doherty, 2012). Recent work suggests that there are at least three behavioral controllers functioning in tandem: habitual, deliberative, and Pavlovian (Daw et al.,
There are two basic possibilities for how preference reversals, and therefore precommitment, arise within the context of these multiple systems. The first possibility is that preference reversals are inherent within a single instrumental system. For example, precommitment may arise entirely within the habitual system as a consequence of multiple exponential discount rates operating in parallel. In this case, precommitment would exist even without an interaction between multiple systems, and would occur without conscious anticipation of a preference reversal; it would occur entirely as a consequence of differential reinforcement (Ainslie,
The second possibility is that preference reversals stem from interactions between systems (Bechara et al.,
In other words, the deliberative system would have insight into the expected future impulsive choice of the habitual system, and would choose to take an action leading to a situation where the habitual or Pavlovian system would not have the impulsive action available. Interestingly, explicit insight or cognitive recognition of future impulsivity is sometimes assumed to be necessary for precommitment (Baumeister et al.,
These two possibilities suggest different ways in which our model of precommitment (Kurth-Nelson and Redish,
Computational Psychiatry
Psychiatry is the study of dysfunction within cognitive and decision-making systems. Whereas traditional psychiatry classifies dysfunctions into categories based on external similarities, new proposals have suggested that classification would be better served by addressing the underlying dysfunction. The emerging field of computational psychiatry suggests that computational models of underlying neural mechanisms can provide a more reasoned basis for the nature of dysfunction and the modality of treatment (Redish et al., 2008; Maia and Frank, 2011; Montague et al., 2012).
Impulsivity is a strong candidate for such a trans-disease mechanism (Bickel et al.,
Precommitment is a powerful strategy to combat impulsivity. Although addicts have faster discounting rates on average than non-addicts (Bickel and Marsch,
Models of precommitment (Kurth-Nelson and Redish,
The latter is particularly interesting in light of the fact that it is possible to change an individual’s discounting function. For example, Bickel et al. (
Finally, the TDRL model depends on having a state-space where precommitment is available as an option. This opens the very important and poorly explored question of how the brain constructs the state-space. In the context of the issues examined here, the brain needs to recognize that precommitment is available. It may be that factors such as working memory and other cognitive resources are important for flexibly constructing adaptive state-spaces, and this may be an essential part of recovery. Even verbally instructing an individual that precommitment is available might be enough to help create the state-space that TDRL or other learning processes could use for precommitment. The ability to form representations of the world that support healthy strategies, even in the face of high underlying impulsivity, may be one of the most important factors in recovery from disorders like addiction.
Statements
Acknowledgments
This work was supported by NIH grant R01 DA024080 (A. David Redish) and by the Max Planck Institute for Human Development as part of the Joint Initiative on Computational Psychiatry and Aging Research between the Max Planck Society and University College London (Zeb Kurth-Nelson). The Wellcome Trust Centre for Neuroimaging is supported by core funding from the Wellcome Trust 091593/Z/10/Z. We thank Warren Bickel for providing data for Figure 3.
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Key Concept
- Precommitment
Taking away a choice from one’s future self in order to enforce one’s present preferences.
- Delay discounting
The attenuation in subjective value of rewards that will be delivered in the future. Delay discounting is typically measured by posing decisions between smaller immediate rewards and larger delayed rewards. Subjects with steep discounting will demand a large increase in the magnitude of a reward in order to tolerate a delay in its receipt.
- Preference reversal
An instability in preferences over time, such that at one time, X is preferred over Y, but at another time, Y is preferred over X. Preference reversal is central to impulsivity disorders. For example, drugs are rarely preferred over healthy choices when the choice is viewed from a distance, but often preferred when immediately available. Therefore it is critical to have mechanisms to enforce the healthy preferences.
- Temporal difference reinforcement learning (TDRL)
A standard computational framework that helps to explain behavioral and neural data. TDRL works by calculating a prediction error at each time step, which encodes the difference between expected and actual reward. This prediction error is used to update expectations such that future prediction errors are minimized. A signal resembling this prediction error is coded by midbrain dopamine neurons.
- Multiple-systems theory of decision-making
Machine learning research shows that there are different computational approaches to solving the problem of producing behavior that maximizes reward. Neural recordings suggest that each of these different algorithms are implemented in the brain, in distinct but overlapping areas.
- Impulsivity
Impulsivity can refer to the inability to inhibit ongoing actions, inability to stick with a long-term plan, or unwillingness to make effort or wait to get a reward. Each of these phenomena reflects a lack of top-down or executive control. In this paper, we focus on unwillingness to wait for delayed rewards.
Zeb L. Kurth-Nelson is a postdoc at the Wellcome Trust Centre for Neuroimaging at University College London. He received his Ph.D. in Neuroscience from the University of Minnesota in 2009. His research interests concern the neural substrates of decision-making and the dysfunction of decision-making in psychiatric disorders such as addiction. z.kurth-nelson@ucl.ac.uk
A. David Redish is currently a professor in the Department of Neuroscience at the University of Minnesota. He has been at the University of Minnesota since 2000, where his lab studies decision-making, particularly issues of covert cognition in rats and failures of decision-making systems in humans.
References
1
AinslieG. (1974). Impulse control in pigeons. J. Exp. Anal. Behav.21, 485–489.10.1901/jeab.1974.21-485
2
AinslieG. (1992). Picoeconomics: The Strategic Interaction of Successive Motivational States Within the Person. New York, NY: Cambridge University Press.
3
AinslieG. (2001). Breakdown of Will. New York, NY: Cambridge University Press.
4
AinslieG.MonterossoJ. R. (2003). Building blocks of self-control: increased tolerance for delay with bundled rewards. J. Exp. Anal. Behav.79, 37–48.10.1901/jeab.2003.79-37
5
AlexanderW. H.BrownJ. W. (2010). Hyperbolically discounted temporal difference learning. Neural. Comput.22, 1511–1527.10.1162/neco.2010.08-09-1080
6
American Psychiatric Association. (2000). Diagnostic and Statistical Manual of Mental Disorders, 4th Edn. Washington, DC: APA.
7
ArielyD.WertenbrochK. (2002). Procrastination, deadlines, and performance: self-control by precommitment. Psychol. Sci.13, 219–224.10.1111/1467-9280.00441
8
BaumeisterR. F.HeathertonT. F.TiceD. M. (1994). Losing Control: How and Why People Fail at Self-Regulation. San Diego, CA: Academic Press.
9
BaumeisterR. F.TierneyJ. (2011). Willpower: Rediscovering the Greatest Human Strength. New York, NY: Penguin Press.
10
BecharaA.NaderK.van der KooyD. (1998). A two-separate-motivational-systems hypothesis of opioid addiction. Pharmacol. Biochem. Behav.59, 1–17.10.1016/S0091-3057(97)00047-6
11
BellmanR. (1957). Dynamic programming. Princeton: Princeton University Press.
12
BickelW. K.JarmolowiczD. P.MuellerE. T.KoffarnusM. N.GatchalianK. M. (2012). Excessive discounting of delayed reinforcers as a trans-disease process contributing to addiction and other disease-related vulnerabilities: emerging evidence. Pharmacol. Ther.134, 287–297.10.1016/j.pharmthera.2012.02.004
13
BickelW. K.MarschL. A. (2001). Toward a behavioral economic understanding of drug dependence: delay discounting processes. Addiction96, 73–86.10.1046/j.1360-0443.2001.961736.x
14
BurksS. V.CarpenterJ. P.GoetteL.RustichiniA. (2009). Cognitive skills affect economic preferences, strategic behavior, and job attachment. Proc. Natl. Acad. Sci. U.S.A. 106, 7745–7750.10.1073/pnas.0812360106
15
ChopraM. P.LandesR. D.GatchalianK. M.JacksonL. C.BickelW. K.BuchhalterA. R.et al (2009). Buprenorphine medication versus voucher contingencies in promoting abstinence from opioids and cocaine. Exp. Clin. Psychopharmacol.17, 226–236.10.1037/a0016597
16
DalleyJ. W.MarA. C.EconomidouD.RobbinsT. W. (2008). Neurobehavioral mechanisms of impulsivity: fronto-striatal systems and functional neurochemistry. Pharmacol. Biochem. Behav.90, 250–260.10.1016/j.pbb.2007.12.021
17
DawN. (2000). Behavioral considerations suggest an average reward TD model of the dopamine system. Neurocomputing32, 679–684.10.1016/S0925-2312(00)00232-0
18
DawN. D. (2003). Reinforcement Learning Models of the Dopamine System and their Behavioral Implications, School of Computer Science, Carnegie Mellon University, Pittsburgh, PA.
19
DawN. D.GershmanS. J.SeymourB.DayanP.DolanR. J. (2011). Model-based influences on humans’ choices and striatal prediction errors. Neuron69, 1204–1215.10.1016/j.neuron.2011.02.027
20
DawN. D.NivY.DayanP. (2005). Uncertainty-based competition between prefrontal and dorsolateral striatal systems for behavioral control. Nat. Neurosci.8, 1704–1711.10.1038/nn1560
21
DayanP.NivY.SeymourB.DawN. D. (2006). The misbehavior of value and the discipline of the will. Neural. Netw.19, 1153–1160.10.1016/j.neunet.2006.03.002
22
FerminA.YoshidaT.ItoM.YoshimotoJ.DoyaK. (2010). Evidence for model-based action planning in a sequential finger movement task. J. Mot. Behav.42, 371–379.10.1080/00222895.2010.526467
23
FrederickS.LoewensteinG.O’DonoghueT. (2002). Time discounting and time preference: a critical review. J. Econ. Lit.40, 351–401.10.1257/002205102320161311
24
GlascherJ.DawN.DayanP.O’DohertyJ. P. (2010). States versus rewards: dissociable neural prediction error signals underlying model-based and model-free reinforcement learning. Neuron66, 585–595.10.1016/j.neuron.2010.04.016
25
GlimcherP. W. (2008). Neuroeconomics: Decision Making and the Brain. San Diego, CA: Academic Press.
26
HaidtJ. (2006). The Happiness Hypothesis: Finding Modern Truth in Ancient Wisdom. New York: Basic Books.
27
HeymanG. M. (2009). Addiction: A Disorder of Choice. Cambridge, MA: Harvard University Press.
28
HuysQ. J.EshelN.O’LionsE.SheridanL.DayanP.RoiserJ. P. (2012). Bonsai trees in your head: how the Pavlovian system sculpts goal-directed choices by pruning decision trees. PLoS Comput. Biol.8, e1002410.10.1371/journal.pcbi.1002410
29
JimuraK.MyersonJ.HilgardJ.KeighleyJ.BraverT. S.GreenL. (2011). Domain independence and stability in young and older adults’ discounting of delayed rewards. Behav. Processes87, 253–259.10.1016/j.beproc.2011.04.006
30
KableJ. W.GlimcherP. W. (2009). The neurobiology of decision: consensus and controversy. Neuron63, 733–745.10.1016/j.neuron.2009.09.003
31
KirbyK. N. (2009). One-year temporal stability of delay-discount rates. Psychon. Bull. Rev.16, 457–462.10.3758/PBR.16.3.457
32
Kurth-NelsonZ.BickelW.RedishA. D. (2012). A theoretical account of cognitive effects in delay discounting. Eur. J. Neurosci.35, 1052–1064.10.1111/j.1460-9568.2012.08058.x
33
Kurth-NelsonZ.RedishA. D. (2009). Temporal-difference reinforcement learning with distributed representations. PLoS ONE.4, e7362.10.1371/journal.pone.0007362
34
Kurth-NelsonZ.RedishA. D. (2010). A reinforcement learning model of precommitment in decision making. Front. Behav. Neurosci.4:184.10.3389/fnbeh.2010.00184
35
Kurth-NelsonZ.RedishA. D. (2012). “Modeling decision-making systems in addiction,” in Computational Neuroscience of Drug Addiction, ed. GutkinB.AhmedS. (New York, NY: Springer), 163–188.
36
KurzbanR. (2010). Why Everyone (else) is A Hypocrite: Evolution and The Modular Mind. Princeton, NJ: Princeton University Press.
37
LadouceurR.BlaszczynskiA.LalandeD. R. (2012). Pre-commitment in gambling: a review of the empirical evidence. Int. Gambl. Stud.1–16.10.1080/14459795.2012.658078
38
MaddenG. J.BickelW. K. (2010). Impulsivity: The Behavioral and Neurological Science of Discounting. Washington, DC: American Psychological Association.
39
MaiaT. V.FrankM. J. (2011). From reinforcement learning models to psychiatric and neurological disorders. Nat. Neurosci.14, 154–162.10.1038/nn.2723
40
MazurJ. E. (1997). Choice, delay, probability, and conditioned reinforcement. Anim. Learn. Behav.25, 131.10.3758/BF03199051
41
McClureS. M.LaibsonD. I.LoewensteinG.CohenJ. D. (2004). Separate neural systems value immediate and delayed monetary rewards. Science306, 503–507.10.1126/science.1100907
42
MontagueP. R.DayanP.SejnowskiT. J. (1996). A framework for mesencephalic dopamine systems based on predictive Hebbian learning. J. Neurosci.16, 1936–1947.
43
MontagueP. R.DolanR. J.FristonK. J.DayanP. (2012). Computational psychiatry. Trends Cogn. Sci. (Regul. Ed.)16, 72–80.10.1016/j.tics.2012.04.003
44
NivY.JoelD.DayanP. (2006). A normative perspective on motivation. Trends Cogn. Sci. (Regul. Ed.)10, 375–381.10.1016/j.tics.2006.06.010
45
O’DohertyJ. P. (2012). Beyond simple reinforcement learning: the computational neurobiology of reward-learning and valuation. Eur. J. Neurosci.35, 987–990.10.1111/j.1460-9568.2012.08074.x
46
OhmuraY.TakahashiT.KitamuraN.WehrP. (2006). Three-month stability of delay and probability discounting measures. Exp. Clin. Psychopharmacol.14, 318–328.10.1037/1064-1297.14.3.318
47
PetryN. M. (2012). Contingency Management for Substance Abuse Treatment: A Guide to Implementing this Evidence-Based Practice. New York: Routledge.
48
RachlinH.GreenL. (1972). Commitment, choice and self-control. J. Exp. Anal. Behav.17, 15–22.10.1901/jeab.1972.17-147
49
RangelA.CamererC.MontagueP. R. (2008). A framework for studying the neurobiology of value-based decision making. Nat. Rev. Neurosci.9, 545–556.10.1038/nrn2357
50
RedishA. D.JensenS.JohnsonA. (2008). A unified framework for addiction: vulnerabilities in the decision process. Behav. Brain Sci.31, 415–437.10.1017/S0140525X0800472X
51
RickS.LoewensteinG. (2008). Intangibility in intertemporal choice. Phil. Trans. Roy. Soc. B363, 3813–3824.10.1098/rstb.2008.0150
52
RobbinsT. W.GillanC. M.SmithD. G.de WitS.ErscheK. D. (2012). Neurocognitive endophenotypes of impulsivity and compulsivity: towards dimensional psychiatry. Trends Cogn. Sci. (Regul. Ed.)16, 81–91.10.1016/j.tics.2011.11.009
53
RoeschM. R.EsberG. R.LiJ.DawN. D.SchoenbaumG. (2012). Surprise! neural correlates of pearce-hall and rescorla-wagner coexist within the brain. Eur. J. Neurosci.35, 1190–1200.10.1111/j.1460-9568.2011.07986.x
54
RomerD.BetancourtL. M.BrodskyN. L.GiannettaJ. M.YangW.HurtH. (2011). Does adolescent risk taking imply weak executive function? A prospective study of relations between working memory performance, impulsivity, and risk taking in early adolescence. Dev. Sci.14, 1119–1133.10.1111/j.1467-7687.2011.01061.x
55
SchultzW. (1998). Predictive reward signal of dopamine neurons. J. Neurophysiol.80, 1–27.
56
SchultzW.DayanP.MontagueP. R. (1997). A neural substrate of prediction and reward. Science275, 1593–1599.10.1126/science.275.5306.1593
57
SchweighoferN.BertinM.ShishidaK.OkamotoY.TanakaS. C.YamawakiS.DoyaK. (2008). Low-serotonin levels increase delayed reward discounting in humans. J. Neurosci.28, 4528–4532.10.1523/JNEUROSCI.4982-07.2008
58
SimonD. A.DawN. D. (2011). Neural correlates of forward planning in a spatial decision task in humans. J. Neurosci.31, 5526–5539.10.1523/JNEUROSCI.3772-11.2011
59
StrotzR. H. (1955). Myopia and inconsistency in dynamic utility maximization. Rev. Econ. Stud.23, 165–180.10.2307/2295722
60
SuttonR. S.BartoA. G. (1998). Reinforcement Learning: An Introduction. Cambridge, MA: MIT Press.
61
TakahashiT.FurukawaA.MiyakawaT.MaesatoH.HiguchiS. (2007). Two-month stability of hyperbolic discount rates for delayed monetary gains in abstinent inpatient alcoholics. Neuro Endocrinol. Lett.28, 131–136.
62
TanakaS. C.SchweighoferN.AsahiS.ShishidaK.OkamotoY.YamawakiS.DoyaK. (2007). Serotonin differentially regulates short- and long-term prediction of rewards in the ventral and dorsal striatum. PLoS ONE2, e1333.10.1371/journal.pone.0001333
63
van der MeerM. A. A.Kurth-NelsonZ.RedishA. D. (2012). Information processing in decision-making systems. Neuroscientist18, 342–359.10.1177/1073858411435128
64
VohsK. D.FaberR. J. (2007). Spent resources: self-regulatory resource availability affects impulse buying. J. Consum. Res. 33, 537–547.10.1086/510228
65
VohsK. D.NelsonN. M.BaumeisterR. F.TiceD. M.SchmeichelB. J.TwengeJ. M. (2008). Making choices impairs subsequent self-control: a limited-resource account of decision making, self-regulation, and active initiative. J. Pers. Soc. Psychol.94, 883–898.10.1037/0022-3514.94.5.883
66
WohlM. J. A.LyonM.DonnellyC. L.YoungM. M.MathesonK.AnismanH. (2008). Episodic cessation of gambling: a numerically aided phenomenological assessment of why gamblers stop playing in a given session. Int. Gambl. Stud. 8, 249–263.10.1080/14459790802405855
67
WunderlichK.DayanP.DolanR. J. (2012). Mapping value based planning and extensively trained choice in the human brain. Nat. Neurosci.15, 786–791.10.1038/nn.3068
Summary
Keywords
discounting function, decision-making, neuroeconomics, temporal diference reinforcement learning, precommitment
Citation
Kurth-Nelson Z and Redish AD (2012) Don’t Let Me Do That! – Models of Precommitment. Front. Neurosci. 6:138. doi: 10.3389/fnins.2012.00138
Received
25 June 2012
Accepted
04 September 2012
Published
08 October 2012
Volume
6 - 2012
Edited by
Daeyeol Lee, Yale University School of Medicine, USA
Reviewed by
Christian C. Luhmann, Stony Brook University, USA; Xinying Cai, Washington University in St Louis, USA
Copyright
© 2012 Kurth-Nelson and Redish.
This is an open-access article distributed under the terms of the Creative Commons Attribution License, which permits use, distribution and reproduction in other forums, provided the original authors and source are credited and subject to any copyright notices concerning any third-party graphics etc.
*Correspondence: redish@umn.edu
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.