Abstract
Serotonin (5-HT) is a neuromodulator that has been attributed to cost assessment and harm aversion. In this review, we look at the role 5-HT plays in making decisions when subjects are faced with potential harmful or costly outcomes. We review approaches for examining the serotonergic system in decision-making. We introduce our group’s paradigm used to investigate how 5-HT affects decision-making. In particular, our paradigm combines techniques from computational neuroscience, socioeconomic game theory, human–robot interaction, and Bayesian statistics. We will highlight key findings from our previous studies utilizing this paradigm, which helped expand our understanding of 5-HT’s effect on decision-making in relation to cost assessment. Lastly, we propose a cyclic multidisciplinary approach that may aid in addressing the complexity of exploring 5-HT and decision-making by iteratively updating our assumptions and models of the serotonergic system through exhaustive experimentation.
INTRODUCTION
Recent theoretical work has implicated serotonergic (5-HT) function in critical dimensions of reward versus punishment and invigoration versus inhibition (). Both of these dimensions have influence on a broad range of decision-making elements, including reward processing, impulsivity, reward discounting, predicting punishment, harm aversion, opponency with other neuromodulators, and anxious states (; ; ). In this review, we first provide evidence from the literature indicating several of the proposed functions attributed to the serotonergic system. Next, we discuss approaches that utilized game theory and other behavioral measures along with some metric of serotonergic function. Finally, we introduce a multidisciplinary experimental paradigm to model the role 5-HT plays in decision-making, and present some of our work that has utilized this paradigm.
Our paradigm begins with base assumptions regarding the role of serotonin or other neuromodulators to construct a simulated agent that can adapt to environmental challenges. That adaptive agent, which is either embodied in a robotic platform for human–robot interaction studies or embedded in a computer interface, is incorporated into a game theoretic environment. The data collected from these experiments are then analyzed to support or reject hypotheses about the roles of neuromodulators in specific cognitive functions such as decision-making, which may lead to the use of more sophisticated adaptive agents in subsequent studies.
FUNCTIONAL ROLES OF SEROTONIN
It has been suggested that serotonin influences a broad range of decision-based functions such as reward assessment, cost assessment, impulsivity, harm aversion, and anxious states. This section discusses recent evidence demonstrating the role serotonin has on these decision-based functions.
Though reward processing is a function that has primarily been attributed to the dopaminergic (DA) system, 5-HT has also been associated with reward-related behavior (, ; ; ; ; ; ). Recent single-unit recordings of serotonergic neurons in the monkey dorsal raphe nucleus (DRN), which is a major source of serotonergic innervation in the central nervous system, demonstrated that many of these neurons represent reward information (; ; ). showed that during a saccade task, after target onset but before reward delivery, the activity of many DRN neurons was modulated by the expected reward size. showed that a group of DRN neurons tracked progress toward future delayed reward after the initiation of a saccade and after the value of the trial was revealed. These studies suggest that DRN neurons, which include 5-HT neurons, may influence behavior based on the amount of delay before reward delivery and the value of the reward in future motivational outcomes (; ).
Other anatomical evidence has shown that projections from DRN to reward-related DA regions support 5-HT’s role in both reward and punishment (). A theoretical review by suggested that the 5-HT and dopamine systems primarily activate in opposition and at times in collaboration for goal directed actions. A review by also highlighted possible computational factors of decision-making in brain regions innervated by serotonin and dopamine (for a schematic of the potential interplay between 5-HT and other brain structures, see , Figure 3). 5-HT projections to dopamine areas have been shown to regulate threat avoidance (; ), and an impairment in these projections can lead to impulsivity and addiction (). Altogether, the interaction between these systems allows 5-HT to play various functional roles in decision-making where reward and punishment, as well as invigoration and inhibition, are in opposition.
In addition to reward processing, several studies have investigated serotonin’s involvement in reward and impulsivity by manipulating levels of central 5-HT in humans using the acute tryptophan depletion (ATD) procedure. ATD is a dietary reduction of tryptophan, an amino acid precursor of 5-HT, which causes a rapid decrease in the synthesis and release of the human brain’s central 5-HT, thus affecting behavioral control (). Altering 5-HT levels via ATD influences a subject’s ability to resist a small immediate reward over a larger delayed reward (delay reward discounting; , ; ). As such, subjects that underwent ATD had both an attenuated assessment of delayed reward and a bias toward small reward, which were indicative of impulsive behavior.
Besides reward, 5-HT has also been linked to predicting punishment or harm aversion (; , ; ; ). paired the ATD procedure with a reversal-learning task, demonstrating that subjects under ATD made more prediction errors for punishment-associated stimuli than for reward-associated stimuli. In a related study, utilized the ATD procedure with a Go/No-Go task to show that lowering 5-HT levels resulted in a decrease in punishment-induced inhibition. In a follow up study, they investigated the mechanisms through which 5-HT regulated punishment-induced inhibition by using the ATD procedure paired with their Reinforced Categorization task, a variation on the Go/No-Go task (). Subjects with lowered 5-HT were faster in responding to stimuli predictive of punishments (), indicating a manipulation of some punishment-predicting mechanism associated with standard serotonergic function. Together, these results suggest that 5-HT influences the ability to inhibit actions that predict punishment and to avoid harmful circumstances.
Beyond punishment, 5-HT has been implicated in stress and anxiety (; ). A recent review by proposed a mechanistic model between environmental impact factors and genetic variation of the serotonin transporter (5-HTTLPR), linking to the risk of depression in humans. They argued that genetic variation may be linked to a balance in the brain’s circuitry underlying stressor reactivity and emotion regulation triggered by a stressful event, ultimately leading to depression (). A review by described studies showing that 5-HT function has been tied to an organism’s anxious states triggered by conditioned or unconditioned fear. Together, this work suggests a functional role for 5-HT in the control of anxious states.
In summary, these studies reveal serotonergic modulation of a wide range of decision-based functions including but not limited to reward processing, motivational encoding, punishment prediction, discounting, impulsivity, harm aversion, and anxious states. Building on this body of work, many researchers in the field have utilized their own approaches in studies to better understand the function of serotonin in behavior. In the present paper, we introduce a novel, multi-disciplinary approach to study serotonin’s influence in decision-making that may highlight many of the functions described above. Our paradigm combines techniques from computational neuroscience, socioeconomic game theory, human–robot interaction, and Bayesian statistics.
INVESTIGATION OF DECISION-MAKING USING GAME THEORY AND SEROTONERGIC MANIPULATIONS
Game theory is a toolbox that is utilized in a multitude of disciplines for its ability to quantitatively measure and predict behavior in situations of cooperation and competition (; ; ). It operates on the principle that organisms will balance reward with effort while acting in self-interest to obtain the optimal result in a given situation. Game theory is especially valuable as a venue for studying human behavior because it provides a replicable, predictable, and controlled environment with clearly defined boundaries. These elements are essential when introducing computer agents as opponents.
Game theory has been combined with manipulations of serotonin to help understand its role in socioeconomic decision-making. For example, in the Prisoner’s Dilemma, where subjects either cooperate or defect in a risky situation, it has been shown that ATD increases the prevalence of defecting, which might be considered an impulsive, risk-taking choice (). Similarly, the Ultimatum game is a test of cooperation in which a proposer offers a share of a resource to a receiver, and the receiver can either accept or reject this offer (; ). In studies conducted by incorporating the Ultimatum game with serotonergic manipulations, it was found that subjects under ATD rejected a significantly higher proportion of unfair offers and that decreased serotonin levels correlated with increased dorsal striatal activity induced by costly punishment (). In contrast, subjects that ingested citalopram, an SSRI, were less likely to punish unfairness in the Ultimatum game (). Together, these studies implicate the involvement of the serotonergic system with cost in decision-making, an important result in understanding the cost and reward mechanisms in the brain.
Another notable game that focuses on the investigation of cooperation and social contracts is the Stag Hunt. In the Stag Hunt, two players must independently choose to hunt a high payoff stag cooperatively or a low payoff hare individually. The risk in decision-making lies in the case when only one player chooses stag, resulting in no payoff for that player (). The body of work involving Stag Hunt largely involves simulations with set-strategy agents or human players as opponents (; ; ). More recently, the use of adaptive agents, computer players that learn in real-time, have been gaining popularity in the field of social decision-making (, ). conducted a study in which adaptive agents played a spatiotemporal version of the Stag Hunt game against human subjects in an fMRI scanner, implicating both rostral medial prefrontal cortex and dorsolateral prefrontal cortex in processing uncertainty and sophistication of agent strategy, respectively. Utilizing adaptive agents allows for a dynamic yet controlled behavioral manipulation in subjects, which is useful within game environments, particularly if applied to studying the cost and reward mechanisms of the brain.
Pairing a decision-making task with ATD in the absence of game theory has further illuminated serotonin’s involvement in behavior. This combination has revealed serotonin’s involvement in the reflexive avoidance of relatively immediate small costs in favor of larger future costs with an Information Sampling Task (). The pairing of ATD with a “four-armed bandit” task showed that depleted subjects tended to be both more perseverative and less receptive to reward (). These results show that the combination of a decision-making task with serotonergic manipulation (e.g., ATD) can provide important information about the role serotonin has in the decision-making process. In general, the combination of ATD with a decision-making task provides a useful venue for the exploration of social behavior and the neural correlates of cost and reward in decision-making.
In addition to altering decision-making, reduced 5-HT levels via ATD have been correlated with individual differences in subject behavior (; ). Subjects with high neuroticism and low self-directedness personality traits have been shown to be particularly susceptible to central 5-HT depletion, resulting in decreased selection of delayed larger reward over smaller immediate reward when performing a delayed reward choice task (). Similarly, subjects with low baseline aggression have displayed reduced reactive aggression when performing a competitive reaction time task with depleted 5-HT levels (). The results from these studies provide evidence for individual behavioral differences correlated with central 5-HT manipulation, which may serve as a direction for future study.
In summary, due to the complex nature of the serotonergic system, researchers have utilized several complementary methods to investigate the varied aspects of its behavioral influence. This review introduces a multidisciplinary experimental paradigm to model the role 5-HT plays in decision-making.
A MULTIDISCIPLINARY PARADIGM TO INVESTIGATE THE SEROTONERGIC SYSTEM
Our paradigm combines socioeconomic game theory with embodied models of learning and adaptive behavior (Figure 1A). In particular, we constructed our computational models to reflect 5-HT’s potential interplay with the expected cost of a decision (; ), under the assumption that 5-HT, released by the DRN, can act as an opponent to dopamine. In this case, activation of the 5-HT system may cause an organism to be withdrawn or risk-averse, and the DA system causes the organism to be uninhibited or risk taking (). Within the context of this paper, cost is defined as either the perceived loss of an expected payoff or harm from a potential threat, depending on the scope of the study it is used in. We will compare our present results using this paradigm with other studies, and discuss future steps that may lead to more accurate modeling of serotonin’s proposed role in assessing the tradeoff between cooperation and competition (Figure 1B).
FIGURE 1
PARADIGM OVERVIEW
When considering the complex and highly varied behavior in decision-making during socioeconomic games (
Embodied models have been shown to elicit strong reactions in humans (
In order to further elucidate the role of serotonin and dopamine in decision-making, we have developed a multidisciplinary paradigm that incorporates embodied adaptive agents into interactive game environments (Figure 1A). Our general paradigm includes several key aspects, which we describe in detail below. In brief, we begin with base assumptions founded on previous studies that are used to construct an adaptive agent. That model, alongside set-strategy agents used in control conditions, are either embodied in a robotic platform (Agents: Embodied) or embedded in a computer interface (Agents: Simulated). Those agents are incorporated into a game theoretic environment in both human subject and simulation experiments. Human subject experiments include manipulation with ATD (Human Experiments: ATD). The data collected from these experiments are analyzed to either support or reject specific hypotheses about the role of serotonin in decision-making, or to create new models that explore the theories that emerge from the data.
BASE ASSUMPTIONS
To start, we assume that serotonergic activity in the raphe nucleus is related to the expected cost of a decision. In this case, cost assessment can be related to harm or loss aversion (
Although controversial compared to other neuromodulators, evidence suggests that serotonergic neuromodulation features both tonic and phasic modes of activity (
The association between tonic/phasic neuromodulation and explore/exploit behavior was originally put forth by
These base assumptions have led us to develop models balancing cost and reward in decision-making through simulation of the neuromodulatory systems, which reflects neural activity in the brain and its resulting explorative and exploitive behaviors.
ADAPTIVE AGENT MODELS
Given our base assumptions, we developed adaptive neural models capable of shaping action selection involved in decision-making (Figure 2). In general, these models made decisions based on their assessment of the expected cost and reward of actions, where cost was related to harm or loss aversion (
FIGURE 2

Adaptive agent architectures. (A) General neural network architecture for Hawk-Dove and Chicken studies. The thick arrows represent all-to-all connections. The dotted arrows with the shaded oval represent modulatory plastic connections. Within the Action Neurons region, neurons with excitatory reciprocal connections are represented as arrow-ended lines, and neurons with reciprocal inhibitory connections are represented as dot-ended lines overlaid by a shaded oval, which denotes plasticity. (B) Actor-Critic schematic. The behavior of the adaptive agent used in the Stag Hunt experiment (
Neural network model
Our neural network model, which was used in Hawk-Dove and Chicken human robot interaction studies, simulated neuromodulation and plasticity based on environmental conditions, as well as previous experiences with cost and reward (
The equation for activity of each of the Game-Dependent Input Neurons (ni) were computed as follows:
where b was a constant value dependent on the game played (b = 0.75 for the Hawk-Dove and b = 0.45 for the Chicken games described below), and noise represented neural noise, which was a random number between 0 and 0.25 drawn from a uniform distribution.
The neural activities for the action and neuromodulatory neurons were simulated by a mean firing rate neuron model, where the firing rate of each neuron ranged from 0 (quiescent) to 1 (maximal firing) on a continuous scale. The activity of both Action Neurons was based on their previous firing rates, plastic extrinsic excitatory input from the Game-Dependent Input Neurons, non-plastic intrinsic excitatory input from the opposing action neuron, and non-plastic intrinsic inhibitory input from the opposing action neuron (Figure 2A). In contrast, the activity of both Neuromodulatory Neurons was based on plastic extrinsic excitatory input from the Game-Dependent Input Neurons and previous cost/reward information reflected in the respective firing rates at the previous time step. The equation for the mean firing rate neuron model was:
where t was the current time step, si was the activation level of neuron i, ρiwas a constant set to 0.1 denoting the persistence of the neuron, and Ii was the synaptic input. The synaptic input of the neuron was based on pre-synaptic neural activity, the connection strength of the synapse, and the amount of neuromodulatory activity:
where wij was the synaptic weight from neuron j to neuron i, and nm was the level of neuromodulation, which was the combined average activity of the Cost and Reward neurons. The noise term represented neural input noise and was a random number between -0.5 and 0, drawn from a uniform distribution.
Phasic neuromodulation can have a strong effect on action selection and learning (
After the neural activities for the Action and Neuromodulatory Neurons were computed, a learning rule was applied to the plastic connections (projections from Game-Dependent Input Neurons) of the neural model. The learning rule depended on the current activity of the pre-synaptic neuron, the post-synaptic neuron, the overall activity of the modulatory neurons, and the cost/reward outcome from the game played:
where sj was the pre-synaptic neuron activity level, si was the post-synaptic neuron activity level, nm was the average activity of the Neuromodulatory Neurons, and R was the level of reinforcement based on payoff and cost (Equation 5). The pre-synaptic neuron (sj) in Equation 4 was the most active Game-Dependent Input Neuron (Equation 1). The post-synaptic neuron (si) could be the most active Action neuron, the Cost neuron, or the Reward neuron. The level of reinforcement was given by:

where the RewardReceived and CostReceived were values determined by the positive and negative payoffs, respectively. The values were determined by a payoff matrix specific to the game being played (
Actor-critic model
In addition to the neural network described above, we have implemented more abstract adaptive agents based on our assumptions (Figure 2B). For example, a variation of the Actor-Critic model was used to simulate reward and cost assessment in a Stag-Hunt game (
where r(t) was either the reward or cost at time t, V(s, t) was the Critic’s weight at state s, at time t, and V(s, t–1) was the Critic’s weight for the previous timestep. More specifically, the reward r(t) value corresponds to the agent’s expected value of their selected choice of rewarding stimuli at that timestep, and the cost r(t) is the negative of that value in the case that the expected reward was not fulfilled (i.e., perceived loss). However, other interpretations of cost are possible depending on the game being played. The delta value of Equation 6 was used to update the weights in the Reward and Cost critic tables at every timestep according to the following function:
The Actor weights were the likelihood to execute a particular action at a state and were updated by using the reward and cost information for that state. In the case that the model decided on choice 1 of two choices, the Actor weights were updated based on the following equation:
V(c1, s, t) was the Actor’s state table value for deciding on choice 1 (of two possible actions) in state s at time t. Likewise, V(c2, s, t) was the Actor’s state table value for choice 2 in state s at time t. δ (t) was the delta value from both the Reward and Cost Critics. Thus, the Actor was updated based on the assessment of both the Cost and Reward Critics.The probabilities for selecting either choice 1 or choice 2 were decided using a SoftMax function:
This implementation of the Actor Critic provided a cost-reward tradeoff mechanism for decision-making in game environments, analogous to the interplay between the DA and serotonergic neuromodulatory systems.
GAME ENVIRONMENTS
In our experiments, we utilized both an adaptive neural network (Figure 2A) and an instantiation of the Actor-Critic model (Figure 2B) to investigate cost and reward in games of decision-making.
The adaptive neural network of Figure 2A, coupled with set-strategy models as controls, were both experimentally embodied as robotic agents and embedded in computer simulation within a game theoretic environment to investigate reciprocal social interactions depending on reward and cost assessment. For these experiments, we selected the game of Hawk-Dove, which is similar to the widely studied Prisoner’s Dilemma (
Our version of Hawk-Dove (Figure 3) was played with an adaptive neural network model contesting over a resource with another player in an area referred to as the territory of interest (TOI) (
FIGURE 3

Hawk-Dove game diagram. The game board included a 5 × 5 grid of squares, upon which a territory was marked and the human and neural agent players were placed. The color of the territory reflected the state of the players’ actions. In the Hawk-Dove, two players must compete for a territory, deciding either to be submissive (display) or aggressive (escalate), avoiding or risking injury in hopes of a larger payoff, respectively. © 2012 IEEE. Reprinted, with permission, from
Alongside Hawk-Dove, Chicken (
FIGURE 4

Chicken game diagram. Two toy cars, one driven by the human and one by the neural agent, were placed at opposite ends of a track. The cars started moving toward each other at the same speed and at the same time, at which point players must decide whether to conservatively swerve out of the way, but take a smaller payoff, or take the risk of a collision and continue straight ahead in hopes of a larger payoff. © 2012 IEEE. Reprinted, with permission, from
While games such as the Prisoner’s Dilemma, Hawk-Dove and Chicken are used to explore cost and reward assessment in competitive situations, the socioeconomic game known as the Stag Hunt is better suited to investigate cooperative situations and the formation of social contracts. Evidence suggests that neural responses are different when the social interaction is perceived to be cooperative versus competitive (
FIGURE 5

Stag Hunt game environment. The game board included a 5 × 5 grid of spaces upon which the player (stick figure image), agent (robot image), stag (stag image), and hare (hare image) tokens resided. The screen included a button to start the experiment, the subject’s score for the round, the subject’s overall score for the experiment, the game number, a countdown to the start of the game, and a counter monitoring the game’s timeout. In the game of Stag Hunt, two players attempt to hunt a low-payoff hare alone, or attempt to cooperate with the other player to hunt a large payoff stag. © 2013 by Adaptive Behavior. Reprinted by Permission of SAGE from
TESTING THE ADAPTIVE MODELS IN GAME ENVIRONMENTS
Depending on the goal of the experiment in question, simulations can provide significant information about behavior development in an adaptive model. These experiments often consist of exhaustive model testing with various opponents, environmental conditions, and intrinsic model parameters resulting in various behaviors and strategies that the model may exhibit. With our model of cost and reward modulation (see Neural Network Model), we conducted simulation experiments that revealed that the model was capable of predicting upcoming costs and rewards (
Following simulation, human subject experiments were performed to test the adaptive model’s performance against human players, as well as the subjects’ reactions to playing against both set-strategy and adaptive agents, and the influence of embodied agents on game play. Our first set of human subjects experiments involved ATD, the dietary manipulation described above that temporarily lowers serotonin levels in the central nervous system, resulting in decreased cooperation and lowered harm-aversion (
In our next set of human subject experiments, the participants played the Stag Hunt game with various set-strategy and adaptive simulated agents. Subjects played games against each of five computer strategies, including an adaptive model. In each game, players navigated the game board on a computer (Figure 5), ending the game when either one of the players successfully captured a hare, or both players worked together to capture a stag. The adaptive agent was an instantiation of the Actor-Critic model that weighed cost and reward to make decisions in the environment in a manner much like the serotonergic and DA systems are thought to act in humans (
Altogether, our multidisciplinary paradigm is one of many that are currently utilized in this field to explore the theorized role of the serotonergic system on behavior as related to cost assessment. The results from using this paradigm provide a balanced and informative procedure that incorporates both neuromodulation and behavior with current methods and technology, as described below.
RESULTS OF OUR STUDIES CONDUCTED USING THIS MULTIDISCIPLINARY APPROACH
ADAPTIVE NEURAL NETWORK PLAYING THE HAWK-DOVE GAME
We explored the research question of how the interplay between cost and reward would lead to appropriate decision-making under varying conditions in a game theoretic environment. To test this question, we modeled several predictions as to how the activity of a cost function leads to appropriate action selection in competitive and cooperative environments (
ATD AND EMBODIMENT IN HAWK-DOVE AND CHICKEN GAMES
To test the influence of embodiment and serotonin on decisions where there is a tradeoff between cooperation and competition, we conducted a study that included both embodied and simulated versions of adaptive agents along with manipulation of serotonin in human subjects. We used ATD to reveal the ways humans interacted with these agents in competitive situations via the Hawk-Dove and Chicken games (
Although the small subject size (n = 8) may have contributed to the lack of significant differences in our measurements of both tryptophan-depletion vs. control conditions and embodied agent vs. simulation conditions, there is the possibility that differences between the conditions were masked by subgroups of subjects responding differently across conditions.
COGNITIVE MODELING
To better understand our results at the individual subject level, we implemented a cognitive model to investigate potential behavioral differences in the subjects’ decision-making by examining their propensity to choose the aggressive action (escalate) in the Hawk-Dove game under the various conditions. Because these cognitive models use Bayesian inference to predict subject behavior based on many individual decisions, their predictions were not weakened by a small sample size.
To investigate how ATD and embodiment affected subjects’ decision-making in our previous work (see ATD and Embodiment in Hawk-Dove and Chicken Games), we implemented a cognitive model using hierarchical Bayesian inference. Hierarchical Bayesian inference has been shown to be a highly customizable and reliable way of exploring models of cognitive processes (
We used a hierarchical latent mixture model with Bayesian inference to analyze the individual differences in decision-making arising from alterations in serotonin levels and of agent embodiment (
We showed that subjects separated into two distinct subgroups for the probability to choose the aggressive action (escalate) across the conditions (Figure 6). Our justification for this conclusion was based on the assumption that the effect of ATD/embodiment could vary across individuals as is reinforced by recent evidence suggesting that the effects of ATD give rise to individual differences across subjects (
FIGURE 6

Estimated group identities based on cognitive modeling results. Both plots show each subject’s likelihood to choose the aggressive action (escalate) for the two different conditions. Red and green dots correspond to subjects that showed a respective increased or decreased probability to escalate from their baselines. Error bars show the 95% Bayesian confidence interval of the posterior mean. The x-axes indicate subject numbers, which correspond to the same subjects in the two plots. The y-axes show the Bayesian model’s mean output indicating group affiliation with respect to the subject’s likelihood to escalate, relative to their independently determined baselines. The y-axis value of 1 indicates a strong likelihood of decreasing choices to escalate relative to their baseline level for the conditions, whereas the value of 2 indicates a strong likelihood of increasing choices to escalate relative to their baseline level for the conditions. The group identities were estimated based on: (A) the influence tryptophan depletion had on subjects’ choices for aggressive actions (Escalation, Tryptophan), and (B) the influence an embodied agent had on subjects’ choices for aggressive actions (Escalation, Robot). Cognitive Science Conference and published in the Proceedings (COGSCI 2012, Sapporo, JP) from
To give a full account of the data, the hierarchical model was designed to address individual differences at two levels: the baseline level, which depends on the subjects inherent tendencies, and the additive level, which depends on the interaction between subjects natural tendencies and experimental conditions. In contrast to the results from our population analysis (see ATD and Embodiment in Hawk-Dove and Chicken Games), we found that clustering subjects into two opposing subgroups better represented the data. That is, one group of subjects, in concordance with expectation, had a higher probability to escalate under tryptophan depletion, but another had a lower probability to escalate in the tryptophan-depleted condition (Figure 6). Similarly, we found that two subgroups better predicted the rate of escalation when comparing responses to a robot versus responses to a computer simulation (Figure 6). The formation of these subgroups is not accounted for by the variance of the data in the population analysis (
HUMANS PLAYING STAG HUNT WITH SIMULATED ADAPTIVE AGENTS
In our recent study using the Stag Hunt game, we investigated the variance in behavior of human subjects while playing Stag Hunt against adaptive (cost/reward learning) and set-strategy agents, with the intent of finding a stronger response evoked by adaptive over set-strategy. We found that adaptive agents, controlled by an Actor-Critic model (see Figure 2B and Adaptive Agent Models), caused subjects to invest more time and effort into game play than set-strategy agents (
Similar to our findings with the Hawk-Dove game, the Stag Hunt study also highlighted subject variation when playing games of decision-making. When assessing the ratio of stag-to-hare captures, playing against an adaptive agent appeared to evoke different equilibriums of hunt decisions in individual subjects. It appears that, much like the Hawk-Dove results (
FUTURE DIRECTIONS USING A CYCLIC, MULTIDISCIPLINARY PARADIGM
In an attempt to more accurately model serotonin’s theorized influence on decision-making, we suggest that future experiments improve upon the approach illustrated in Figure 1A with the addition of an iterative component (Figure 1B). Such a cyclic paradigm would have the following components: (1) the development of an embodied neural model to support socioeconomic game studies; (2) an experimental protocol in which subject behavior and neural correlates of decision-making can be probed and categorized; (3) the design of an improved neural model that captures the neuromodulatory influences and individual variation of decision-making in socioeconomic games, to be used in subsequent experiments; and (4) the deployment of a population of models with varying phenotypes to be used in subsequent socioeconomic game studies. Components (3) and (4) allow the paradigm to run cyclically, thereby improving the paradigm through analysis and incorporation of past results. In this cyclic, multidisciplinary paradigm, we amend our previous multidisciplinary paradigm with a feedback loop that: (1) makes new interpretations for the role of serotonin in subject behavior; (2) develops a new cognitive model based on subject behavior; and (3) modifies the adaptive neural network to construct agents that capture individual behavioral differences demonstrated by subjects. These modifications are performed with the intention of refining the adaptive neural network’s performance, making its behavior more natural and human-like. After each cycle, the experimental paradigm improves to better suit the purposes of the task (e.g., stronger decision-making/modeling of neuromodulation), while holding constant the general framework of testing (e.g., game theoretic environment, simulation/human experimentation, etc.).
While the proposed paradigm is intended to improve the field of modeling decision neuroscience, our current models are rather abstract and would benefit from the incorporation of additional empirical data collected from the mammalian brain in neuroimaging and neurophysiological studies. Functional data from neuroimaging studies can help our models become more biologically realistic by revealing the specific brain areas active during select behaviors. Single unit recording studies in animals dictate the more granular neural behavior within each modeled brain region. Together, empirical data provides the base assumptions that guide our models computational neural behavior and architecture, making them more biologically realistic. Improved biological plausibility can, in turn, increase the efficacy of theoretical predictions made by our models, resulting in better theories to be tested through future neurophysiological and neuroimaging experiments.
Single unit recording studies in animals provide a critical component to computational modeling, as physical data is essential for developing base assumptions and confirming predictions made by models. For instance, phasic and tonic serotonergic activity in monkeys and rats has been associated with components of cost and reward processing (
Empirical data from neuroimaging studies provide a relationship between brain activity and behavior that can be used as the foundation for biological plausibility in a computational model. For example, fMRI has been used to determine the relationship between brain regions innervated by serotonin and behaviors involved with reward prediction (
In addition to these sources of empirical evidence, theoretical data from other biologically realistic models and neurally inspired robotic agents can contribute to the biological plausibility of our models and serve as a basis for further empirical investigation. For example, incorporating more biophysically detailed models of DA and serotonergic neuromodulation, such as (
Embodiment is a key element to the paradigm we are promoting, and these human robot interaction experiments may not only evoke strong responses in subjects, but they may also inform the development of future neurorobots. Embodied models using robotic platforms have provided clues as to how neuromodulation can give rise to adaptive behavior in biological systems (
It is important to emphasize that the models used in our paradigm serve as a venue for investigating the influence of serotonin in motivational systems for robots and other autonomous systems. Future iterations of research through our paradigm could modify our model to increase the accuracy and scope of its biological representation. In contrast to other similar models that associate serotonin with decision-making (
Utilizing a cyclic paradigm lends itself especially well to studies that incorporate embodiment, as the internal mechanisms governing embodied models are constantly being updated to improve their behavior during interactions with human subjects. The evidence that human subjects are more likely to treat robotic platforms similarly to other humans rather than computer simulations (
While implementing adaptive agents into robotic platforms is a promising venture for future study, past experiments have revealed individual differences between subjects that warrant the investigation of genetic sources. Because the results from our Hawk-Dove and Stag-Hunt experiments showed individual variation in game play, genetic screening for polymorphisms in human subjects could provide a venue for studying serotonin’s role in this variation. Several groups have suggested that individual differences in behavior are influenced by genetic polymorphisms related to serotonin signaling (
Though it is important to utilize new techniques such as genetic screening to better understand the role of serotonin in decision-making, a primary benefit to our paradigm is its incorporation of theoretical predictions from past work into future studies. Previously, we found that the concept of two opposing subgroups (Figure 6) best described the subjects’ behavior in the Hawk-Dove game. This theoretical data could be applied to the next generation of adaptive model (via the iterative component of Figure 1B) through additional assumptions or constraints of serotonergic neuromodulation. These new assumptions lead to better predictions about the diversity in behavior resulting from serotonergic manipulation. As another example, from the Stag Hunt human subject experiment, we discovered the tendency for adaptive agents to move counterintuitively when subject behavior was erratic. In a second iteration of experiments, we could improve upon this model by utilizing a top-down mechanism founded in neuromodulation to converge behavior in the face of seemingly random influence. By using cognitive models, we can create behavioral phenotypes in future models to match the potential individual differences that arise in any given population. These are a couple of examples of how this cyclic paradigm would help our adaptive models and ultimately our understanding of neuromodulatory influence over behavior in decision-making.
In terms of potential clinical application, the proposed paradigm may help illuminate components of brain disorders associated with abnormal serotonergic function. Serotonin has been implicated in a variety of neuropsychiatric conditions including bipolar disorder (
Thus, the cyclic, multidisciplinary paradigm provides a strong approach toward making predictions about the neurobiology that ties serotonin to motivated behavior. As we continue to explore serotonin and its role in decision-making, future studies should consider applying this paradigm in order to accommodate the complex behavior that accompanies the activity of the serotonergic system. Adaptive neural models situated in a game theoretic environment utilized in both human and simulation experiments, accompanied with analysis that leads to an upgraded model for future use, is a strategy that lends itself to the production of valuable research in the fields of neuromodulation, behavior, technology, and neuropsychiatry.
Statements
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
REFERENCES
1
AsherD. E.BartonB.BrewerA. A.KrichmarJ. L. (2012a). Reciprocity and retaliation in social games with adaptive agents.IEEE Trans. Auton. Ment. Dev.4226–238.
2
AsherD. E.ZhangS.ZaldivarA.LeeM. D.KrichmarJ. L. (2012b). “Modeling individual differences in socioeconomic game playing,” inProceedings of the 34th Annual Conference of the Cognitive Science SocietySapporo.
3
AsherD. E.ZaldivarA.KrichmarJ. L. (2010). “Effect of neuromodulation on performance in game playing: a modeling study,” inIEEE 9th International Conference on Development and Learning (Irvine: University of California) 155–160.
4
Aston-JonesG.CohenJ. D. (2005). An integrative theory of locus coeruleus-norepinephrine function: adaptive gain and optimal performance.Annu. Rev. Neurosci.28403–450. 10.1146/annurev.neuro.28.061604.135709
5
AveryM. C.DuttN.KrichmarJ. L. (2013). A large-scale neural network model of the influence of neuromodulatory levels on working memory and behavior.Front. Comput. Neurosci. 7:133.10.3389/fncom.2013.00133
6
BerridgeK. C.RobinsonT. E. (1998). What is the role of dopamine in reward: hedonic impact, reward learning, or incentive salience?Brain Res. Brain Res. Rev.28309–369. 10.1016/S0165-0173(98)00019-8
7
BevilacquaL.GoldmanD. (2011). Genetics of emotion.Trends Cogn. Sci.15401–408. 10.1016/j.tics.2011.07.009
8
Bolado-GomezR.GurneyK. (2013). A biologically plausible embodied model of action discovery.Front. Neurorobot.7:4. 10.3389/fnbot.2013.00004
9
BoureauY. L.DayanP. (2010). Opponency revisited: competition and cooperation between dopamine and serotonin.Neuropsychopharmacology3674–97. 10.1038/npp.2010.151
10
BreazealC.ScassellatiB. (2002). Robots that imitate humans.Trends Cogn. Sci.6481–488. 10.1016/S1364-6613(02)02016-8
11
BriandL. A.GrittonH.HoweW. M.YoungD. A.SarterM. (2007). Modulators in concert for cognition: modulator interactions in the prefrontal cortex.Prog. Neurobiol.8369–91. 10.1016/j.pneurobio.2007.06.007
12
Bromberg-MartinE. S.HikosakaO.NakamuraK. (2010). Coding of task reward value in the dorsal raphe nucleus.J. Neurosci.306262–6272. 10.1523/JNEUROSCI.0015-10.2010
13
Cano-ColinoM.AlmeidaR.CompteA. (2013). Serotonergic modulation of spatial working memory: predictions from a computational network model.Front. Integr. Neurosci.7:71. 10.3389/fnint.2013.00071
14
ChorleyP.SethA. K. (2011). Dopamine-signaled reward predictions generated by competitive excitation and inhibition in a spiking neural network model.Front. Comput. Neurosci.5:21. 10.3389/fncom.2011.00021
15
CoolsR.NakamuraK.DawN. D. (2010). Serotonin and dopamine: unifying affective, activational, and decision functions.Neuropsychopharmacology3698–113. 10.1038/npp.2010.121
16
CoolsR.RobertsA. C.RobbinsT. W. (2008). Serotoninergic regulation of emotional and behavioural control processes.Trends Cogn. Sci.1231–40. 10.1016/j.tics.2007.10.011
17
CoolsR.RobinsonO. J.SahakianB. (2007). Acute tryptophan depletion in healthy volunteers enhances punishment prediction but does not affect reward prediction.Neuropsychopharmacology332291–2299. 10.1038/sj.npp.1301598
18
CoxB.KrichmarJ. (2009). Neuromodulation as a robot controller.IEEE Robot. Autom. Mag.1672–80. 10.1109/MRA.2009.933628
19
CraigA. B.AsherD. E.OrosN.BrewerA. A.KrichmarJ. L. (2013). Social contracts and human–computer interaction with simulated adapting agents.Adapt. Behav.21371–387. 10.1177/1059712313491612
20
CramerJ. S. (2003). “The origins and development of the logit model,” in Logit Models from Economics and Other Fields, Chap.9 (Cambridge: Cambridge UniversityPress)149–158.
21
CrockettM. J.ClarkL.RoiserJ. P.RobinsonO. J.CoolsR.ChaseH. W.et al (2011). Converging evidence for central 5-HT effects in acute tryptophan depletion.Mol. Psychiatry17121–123. 10.1038/mp.2011.106
22
CrockettM. J.Apergis-SchouteA.HerrmannB.LiebermanM.MullerU.RobbinsT. W.et al (2013). Serotonin modulates striatal responses to fairness and retaliation in humans.J. Neurosci.333505–3513. 10.1523/JNEUROSCI.2761-12.2013
23
CrockettM. J.ClarkL.Apergis-SchouteA. M.Morein-ZamirS.RobbinsT. W. (2012). Serotonin modulates the effects of Pavlovian aversive predictions on response vigor.Neuropsychopharmacology372244–2252. Available at: http://www.nature.com/npp/journal/vaop/ncurrent/full/npp201275a.html10.1038/npp.2012.75
24
CrockettM. J.ClarkL.LiebermanM. D.TabibniaG.RobbinsT. W. (2010). Impulsive choice and altruistic punishment are correlated and increase in tandem with serotonin depletion.Emotion10855–862. 10.1037/a0019861
25
CrockettM. J.ClarkL.RobbinsT. W. (2009). Reconciling the role of serotonin in behavioral inhibition and aversion: acute tryptophan depletion abolishes punishment-induced inhibition in humans.J. Neurosci.2911993–11999. 10.1523/JNEUROSCI.2513-09.2009
26
CrockettM. J.ClarkL.TabibniaG.LiebermanM. D.RobbinsT. W. (2008). Serotonin modulates behavioral reactions to unfairness.Science3201739–1739. 10.1126/science.1155577
27
DawN. D.KakadeS.DayanP. (2002). Opponent interactions between serotonin and dopamine.Neural Netw.15603–616. 10.1016/S0893-6080(02)00052-7
28
DayanPHuysQ. J. M. (2009). Serotonin in affective control.Annu. Rev. Neurosci.3295–126. 10.1146/annurev.neuro.051508.135607
29
DeakinJ. F. W. (2003). Depression and antisocial personality disorder: two contrasting disorders of 5HT function.J. Neural Transm. Suppl.79–93.
30
de QuervainD. J.-F. (2004). The neural basis of altruistic punishment.Science3051254–1258. 10.1126/science.1100735
31
DemotoY.OkadaG.OkamotoY.KunisatoY.AoyamaS.OnodaK.et al (2012). Neural and personality correlates of individual differences related to the effects of acute tryptophan depletion on future reward evaluation.Neuropsychobiology6555–64. 10.1159/000328990
32
Di CaraB.DusticierN.ForniC.LievensJ. C.DaszutaA. (2001). Serotonin depletion produces long lasting increase in striatal glutamatergic transmission.J. Neurochem.78240–248. 10.1046/j.1471-4159.2001.00242.x
33
DoyaK. (2002). Metalearning and neuromodulation.Neural Netw.15495–506. 10.1016/S0893-6080(02)00044-8
34
DoyaK. (2008). Modulators of decision making.Nat. Neurosci.11410–416. 10.1038/nn2077
35
FleissbachK.WeberB.TrautnerP.DohmenT.SundeU.ElgerC.et al (2007). Social comparison affects reward-related brain activity in the human ventral striatum.Science3181305–1308. 10.1126/science.1145876
36
FletcherM. L.ChenW. R. (2010). Neural correlates of olfactory learning: critical role of centrifugal neuromodulation.Learn. Mem.17561–570. 10.1101/lm.941510
37
GuQ. (2002). Neuromodulatory transmitter systems in the cortex and their role in cortical plasticity.Neuroscience111815–835. 10.1016/S0306-4522(02)00026-X
38
HasselmoM. E.McGaughyJ. (2004). High acetylcholine levels set circuit dynamics for attention and encoding and low acetylcholine levels set dynamics for consolidation.Prog. Brain Res.145207–231. 10.1016/S0079-6123(03)45015-2
39
HeislerL. K.ChuH.-M.BrennanT. J.DanaoJ. A.BajwaP.ParsonsL. H.et al (1998). Elevated anxiety and antidepressant-like responses in serotonin 5-HT1A receptor mutant mice.Proc. Natl. Acad. Sci. U.S.A.9515049–15054. 10.1073/pnas.95.25.15049
40
HombergJ. R.LeschK.-P. (2011). Looking on the bright side of serotonin transporter gene variation.Biol. Psychiatry69513–519. 10.1016/j.biopsych.2010.09.024
41
HydeL. W.BogdanR.HaririA. R. (2011). Understanding risk for psychopathology through imaging gene–environment interactions.Trends Cogn. Sci.15417–427. 10.1016/j.tics.2011.07.001
42
JasinskaA. J.LowryC. A.BurmeisterM. (2012). Serotonin transporter gene, stress and raphe–raphe interactions: a molecular mechanism of depression.Trends Neurosci.35395–402. 10.1016/j.tins.2012.01.001
43
KiddC. D.BreazealC. (2004). “Effect of a robot on user perceptions,” inIntelligent Robots and Systems, 2004. Proceedings of 2004 IEEE/RSJ International Conference,Vol. 43559–3564. Available at: http://ieeexplore.ieee.org/xpls/abs_all.jsp?arnumber=1389967
44
KieslerS.SproullL.WatersK. (1996). A prisoner’s dilemma experiment on cooperation with people and human-like computers.J. Pers. Soc. Psychol.704710.1037/0022-3514.70.1.47
45
KobayashiM.ImamuraK.SugaiT.OnodaN.YamamotoM.KomaiS.et al (2000). Selective suppression of horizontal propagation in rat visual cortex by norepinephrine.Eur. J. Neurosci.12264–272. 10.1046/j.1460-9568.2000.00917.x
46
KrämerU. M.JansmaH.TempelmannC.MünteT. F. (2007). Tit-for-tat: the neural basis of reactive aggression.Neuroimage38203–211. 10.1016/j.neuroimage.2007.07.029
47
KrämerU. M.RibaJ.RichterS.MünteT. F. (2011). An fMRI study on the role of serotonin in reactive aggression.PLoS ONE6:e27668. 10.1371/journal.pone.0027668
48
KrichmarJ. L. (2008). The neuromodulatory system: a framework for survival and adaptive behavior in a challenging world.Adapt. Behav.16385–399. 10.1177/1059712308095775
49
KrichmarJ. L. (2013). A neurorobotic platform to test the influence of neuromodulatory signaling on anxious and curious behavior.Front. Neurorobot.7:1. 10.3389/fnbot.2013.00001
50
KrichmarJ. L.EdelmanG. M. (2002). Machine psychology: autonomous behavior, perceptual categorization and conditioning in a brain-based device.Cereb. Cortex12818–830. 10.1093/cercor/12.8.818
51
KrichmarJ. L.EdelmanG. M. (2005). Brain-based devices for the study of nervous systems and the development of intelligent machines.Artif. Life1163–77. 10.1162/1064546053278946
52
LapishC. C.KroenerS.DurstewitzD.LavinA.SeamansJ. K. (2006). The ability of the mesocortical dopamine system to operate in distinct temporal modes.Psychopharmacology191609–625. 10.1007/s00213-006-0527-8
53
LeeD. (2008a). Game theory and neural basis of social decision making.Nat. Neurosci.11404–409. 10.1038/nn2065
54
LeeM. D. (2008b). Three case studies in the Bayesian analysis of cognitive models.Psychon. Bull. Rev.151–15. 10.3758/PBR.15.1.1
55
LeeM. D.ZhangS.MunroM.SteyversM. (2011). Psychological models of human and optimal performance in bandit problems.Cogn. Syst. Res.12164–174. 10.1016/j.cogsys.2010.07.007
56
LiC.LoweR.ZiemkeT. (2013). Humanoids learning to walk: a natural CPG-actor-critic architecture.Front. Neurorobot. 7:5.10.3389/fnbot.2013.00005
57
LothE.CarvalhoF.SchumannG. (2011). The contribution of imaging genetics to the development of predictive markers for addictions.Trends Cogn. Sci.15436–446. 10.1016/j.tics.2011.07.008
58
LowryC. A.HaleM. W.EvansA. K.HeerkensJ.StaubD. R.GasserP. J.et al (2008). Serotonergic systems, anxiety, and affective disorder.Ann. N. Y. Acad. Sci.114886–94. 10.1196/annals.1410.004
59
LuciwM.KompellaV.KazerounianS.SchmidhuberJ. (2013). An intrinsic value system for developing multiple invariant representations with incremental slowness learning.Front. Neurorobot. 7:9.10.3389/fnbot.2013.00009
60
Maynard SmithJ. (1982). Evolution and the Theory of Games.Cambridge; New York: Cambridge University Press.
61
McClureS. M.DawN. DRead MontagueP. (2003). A computational substrate for incentive salience.Trends Neurosci.26423–428. 10.1016/S0166-2236(03)00177-2
62
MillanM. J. (2003). The neurobiology and control of anxious states.Prog. Neurobiol.7083–244. 10.1016/S0301-0082(03)00087-X
63
MoranR. J.CampoP.SymmondsM.StephanK. E.DolanR. J.FristonK. J. (2013). Free energy, precision and learning: the role of cholinergic neuromodulation.J. Neurosci.338227–8236. 10.1523/JNEUROSCI.4255-12.2013
64
MurphyS. E.LonghitanoC.AyresR. E.CowenP. J.HarmerC. J.RogersR. D. (2009). The role of serotonin in nonnormative risky choice: the effects of tryptophan supplements on the ’reflection effect’ in healthy adult volunteers.J. Cogn. Neurosci.211709–1719. 10.1162/jocn.2009.21122
65
NakamuraK. (2013). The role of the dorsal raphé nucleus in reward-seeking behavior.Front. Integr. Neurosci.7:60. 10.3389/fnint.2013.00060
66
NakamuraK.MatsumotoM.HikosakaO. (2008). Reward-dependent modulation of neuronal activity in the primate dorsal raphe nucleus.J. Neurosci.285331–5343. 10.1523/JNEUROSCI.0021-08.2008
67
NewellB. R.LeeM. D. (2011). The right tool for the job? Comparing an evidence accumulation and a naive strategy selection model of decision making.J. Behav. Decis. Mak.24456–481. 10.1002/bdm.703
68
NishizawaS.BenkelfatC.YoungS. N.LeytonM.MzengezaS.De MontignyC.et al (1997). Differences between males and females in rates of serotonin synthesis in human brain.Proc. Natl. Acad. Sci. U.S.A.945308–5313. 10.1073/pnas.94.10.5308
69
NowakM. A.PageK. M.SigmundK. (2000). Fairness versus reason in the ultimatum game.Science2891773–1775. 10.1126/science.289.5485.1773
70
OkadaK.NakamuraK.KobayashiY. (2011). A neural correlate of predicted and actual reward-value information in monkey pedunculopontine tegmental and dorsal raphe nucleus during saccade tasks.Neural Plast.20111–21. 10.1155/2011/579840
71
RapaportA.ChammahA. M. (1966). The game of chicken.Am. Behav. Sci.1010–28. 10.1177/000276426601000303
72
RedgraveP.GurneyK. (2006). The short-latency dopamine signal: a role in discovering novel actions?Nat. Rev. Neurosci.7967–975. 10.1038/nrn2022
73
RillingJ. K.SanfeyA. G. (2011). The neuroscience of social decision-making.Annu. Rev. Psychol.6223–48. 10.1146/annurev.psych.121208.131647
74
RobinsonO.CoolsR.CrockettM.SahakianB. (2009). Mood state moderates the role of serotonin in cognitive biases.J. Psychopharmacol.24573–583. 10.1177/0269881108100257
75
RouderJ. N.LuJ.SpeckmanP.SunD.JiangY. (2005). A hierarchical model for estimating response time distributions.Psychon. Bull. Rev.,12195–223. 10.3758/BF03257252
76
RudebeckP. H.WaltonM. E.SmythA. N.BannermanD. MRushworthM. F. S. (2006). Separate neural pathways process different decision costs.Nat. Neurosci.91161–1168. 10.1038/nn1756
77
SanfeyA. G. (2003). The neural basis of economic decision-making in the ultimatum game.Science3001755–1758. 10.1126/science.1082976
78
ScholzJ. T.WhitemanM. A. (2010). Social Capital in Coordination Experiments: Risk, Trust and Position.Available at: http://opensiuc.lib.siu.edu/pn_wp/50/
79
SchultzW. (1997). A neural substrate of prediction and reward.Science2751593–1599. 10.1126/science.275.5306.1593
80
SchweighoferN.BertinM.ShishidaK.OkamotoY.TanakaS. C.YamawakiS.et al (2008). Low-serotonin levels increase delayed reward discounting in humans.J. Neurosci.284528–4532. 10.1523/JNEUROSCI.4982-07.2008
81
SchweimerJ. V.UnglessM. A. (2010). Phasic responses in dorsal raphe serotonin neurons to noxious stimuli.Neuroscience1711209–1215. 10.1016/j.neuroscience.2010.09.058
82
SeymourB.DawN. D.RoiserJ. P.DayanP.DolanR. (2012). Serotonin selectively modulates reward value in human decision-making.J. Neurosci.325833–5842. 10.1523/JNEUROSCI.0053-12.2012
83
ShumakeJ.IlangoA.ScheichH.WetzelW.OhlF. W. (2010). Differential neuromodulation of acquisition and retrieval of avoidance learning by the lateral habenula and ventral tegmental area.J. Neurosci.305876–5883. 10.1523/JNEUROSCI.3604-09.2010
84
SkyrmsB. (2001). The Stag Hunt. Presented at the Presidential Address Pacific Division of the American Philosophical Association.
85
SkyrmsB. (2004). The Stag Hunt and the Evolution of Social Structure.New York: Cambridge University Press.
86
StrobelA.ZimmermannJ.SchmitzA.ReuterM.LisS.WindmannS.et al (2011). Beyond revenge: neural and genetic bases of altruistic punishment.Neuroimage54671–680. 10.1016/j.neuroimage.2010.07.051
87
SzolnokiA.PercM. (2008). Promoting cooperation in social dilemmas via simple coevolutionary rules.Eur. Phys. J. B67337–344. 10.1140/epjb/e2008-00470-8
88
TakahashiH.TakanoH.CamererC. F.IdenoT.OkuboS.MatsuiH.et al (2012). Honesty mediates the relationship between serotonin and reaction to unfairness.Proc. Natl. Acad. Sci. U.S.A.1094281–4284. 10.1073/pnas.1118687109
89
TanakaS. C.SchweighoferN.AsahiS.ShishidaK.OkamotoY.YamawakiS.et al (2007). Serotonin differentially regulates short- and long-term prediction of rewards in the ventral and dorsal striatum.PLoS ONE2:e1333. 10.1371/journal.pone.0001333
90
TanakaS. C.ShishidaK.SchweighoferN.OkamotoY.YamawakiS.DoyaK. (2009). Serotonin affects association of aversive outcomes to past actions.J. Neurosci.2915669–15674. 10.1523/JNEUROSCI.2799-09.2009
91
TopsM.RussoS.BoksemM. A. S.TuckerD. M. (2009). Serotonin: modulator of a drive to withdraw.Brain Cogn.71427–436. 10.1016/j.bandc.2009.03.009
92
ValluriA. (2006). Learning and cooperation in sequential games.Adapt. Behav.14195–209. 10.1177/105971230601400304
93
WainerJ.Feil-SeiferD. J.ShellD. A.MataricM. J. (2007). “Embodiment and human–robot interaction: a task-based perspective,” inRobot and Human interactive Communication, 2007. RO-MAN 2007. The 16th IEEE International Symposium.872–877. Available at: http://ieeexplore.ieee.org/xpls/abs_all.jsp?arnumber = 4415207
94
WengJ.LuciwM. D.ZhangQ. (2013). Brain-like emergent temporal processing: emergent open states.IEEE Trans. Auton. Ment. Dev.589–116. 10.1109/TAMD.2013.2258398
95
WetzelsR.VandekerckhoveJ.TuerlinckxF.WagenmakersE.-J. (2010). Bayesian parameter estimation in the Expectancy Valence model of the Iowa gambling task.J. Math. Psychol.5414–27. 10.1016/j.jmp.2008.12.001
96
Wong-LinK.JoshiA.PrasadG.McGinnityT. M. (2012). Network properties of a computational model of the dorsal raphe nucleus.Neural Netw.3215–25. 10.1016/j.neunet.2012.02.009
97
WoodR. M.RillingJ. K.SanfeyA. G.BhagwagarZ.RogersR. D. (2006). Effects of tryptophan depletion on the performance of an iterated prisoner’s dilemma game in healthy adults.Neuropsychopharmacology311075–1084. 10.1038/sj.npp.1300932
98
YoshidaW.DolanR. J.FristonK. J. (2008). Game theory of mind.PLoS Comput. Biol.4:e1000254. 10.1371/journal.pcbi.1000254
99
YoshidaW.SeymourB.FristonK. J.DolanR. J. (2010). Neural mechanisms of belief inference during cooperative games.J. Neurosci.3010744–10751. 10.1523/JNEUROSCI.5895-09.2010
100
ZaldivarA.AsherD.KrichmarJ. (2010). Simulation of how neuromodulation influences cooperative behavior.Lect. Notes Comput. Sci.11649–660. 10.1007/978-3-642-15193-4_61
Summary
Keywords
serotonin, embodiment, cost assessment, human–robot interaction, adaptive agents, game theory, cognitive modeling, acute tryptophan depletion
Citation
Asher DE, Craig AB, Zaldivar A, Brewer AA and Krichmar JL (2013) A dynamic, embodied paradigm to investigate the role of serotonin in decision-making. Front. Integr. Neurosci. 7:78. doi: 10.3389/fnint.2013.00078
Received
20 March 2013
Accepted
24 October 2013
Published
21 November 2013
Volume
7 - 2013
Edited by
Kae Nakamura, Kansai Medical University, Japan
Reviewed by
Lei Niu, Albert Einstein College of Medicine, USA; KongFatt Wong-Lin, University of Ulster, Northern Ireland
Copyright
© 2013 Asher, Craig, Zaldivar, Brewer and Krichmar.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Derrik E. Asher, Alexis B. Craig, and Andrew Zaldivar, Cognitive Anteater Robotics Lab, Department of Cognitive Sciences, University of California, 2220 Social and Behavioral Sciences Gateway Building, Irvine, CA 92697, USA e-mail: dasher@uci.edu; acraig1@uci.edu; azaldiva@uci.edu
†Derrik E. Asher, Alexis B. Craig, and Andrew Zaldivar have contributed equally to this work.
This article was submitted to the journal Frontiers in Integrative Neuroscience.
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.