SYSTEMATIC REVIEW article

Front. Cognit., 02 January 2026

Sec. Reason and Decision-Making

Volume 4 - 2025 | https://doi.org/10.3389/fcogn.2025.1678665

Music's context-dependent influence on oxytocin, social bonding, and emotion regulation: a systematic review

  • 1. Graduate Institute of Musicology, National Taiwan University, Taipei, Taiwan

  • 2. Graduate Institute of Brain and Mind Sciences, National Taiwan University, Taipei, Taiwan

Abstract

Objective:

This systematic review aims to explore how individual and group musical activities influence social bonding and emotion regulation through the oxytocinergic system.

Methods:

Following the PRISMA 2020 guidelines, a systematic search of PubMed, Embase, Scopus, Web of Science Core Collection, and PsycInfo was conducted to identify studies up to October 2024, supplemented by a manual search. One reviewer screened studies, extracted data, and assessed the study quality. Framework synthesis and narrative synthesis were conducted to integrate findings.

Results:

A total of 1,865 records were identified. After reviewing the full-text papers, 20 studies (seven randomized controlled trials and 13 quasi-experiments) were included, which involved 877 participants across healthy and clinical populations. The reviewed interventions included singing, playing instruments, listening to music, and music therapy. Most studies reported improvements in psychosocial outcomes, such as reduced anxiety and depression or enhanced social cognition, but they do not always align with peripheral oxytocin (OXT) changes. However, certain psychosocial outcomes or contexts revealed relatively consistent patterns in OXT responses, suggesting the presence of context-dependent modulation. Short-term interventions often reported detectable peripheral OXT changes, which only partially reflected the temporary activity of magnocellular OXT neurons in the hypothalamus. No significant changes in baseline peripheral OXT levels were observed after long-term interventions.

Conclusion:

Music-induced OXT responses are context-dependent. The bidirectional modulation of OXT supports social bonding and emotion regulation in musical contexts. Clinicians and music therapists should carefully consider therapeutic goals, individual differences, and environmental factors when designing music therapy.

1 Introduction

There is growing interest in music therapy, which has shown promise in promoting mental health and wellbeing (; ). Recent studies have begun to explore the underlying neural mechanisms of music's effects, with a particular focus on oxytocin (OXT), a crucial neuropeptide for social bonding and emotion regulation ().

OXT is synthesized by the paraventricular nucleus (PVN) and supraoptic nucleus (SON) of the hypothalamus and released into the central nervous system (CNS) and endocrine system to exert its effects (). In the endocrine system, OXT is secreted by the posterior pituitary into the bloodstream to support reproductive processes like uterine contraction and lactation (; ). In the CNS, OXT is released from dendritic or axonal terminals to specific brain regions to modulate neurotransmitter release, synaptic plasticity, and neural network activity (; ). For example, OXT promotes the release of γ-aminobutyric acid (GABA), which reduces neuronal excitability and produces anxiolytic effects (; Yuki, 2023). OXT engages in intracellular signaling to regulate gene expression (; Uvnas-Moberg, 2024) and synaptic plasticity in the amygdala (), hippocampus (), and prefrontal cortex () to support long-term adaptations in emotional responses and social behaviors (). OXT can suppress the expression of corticotropin-releasing hormone (CRH) and reduce downstream cortisol secretion to moderate stress responses (; ; ).

OXT also plays a role in social cognition (). It enhances an individual's ability to recognize, process, and respond to social cues. Notably, these effects are influenced by contextual factors (e.g., safe or threatening environments), and individual factors, including sex, hormonal status, gene variations, attachment style, history of childhood trauma, and the presence of psychiatric disorders, influence the OXTergic system responsiveness (). These findings suggest that the role of OXT is to enhance sensitivity to environmental changes for appropriate behavioral selection ().

It is unclear how music modulates the OXTergic system to influence emotion regulation and social bonding. Although empirical studies have suggested that music can lead to peripheral OXT changes, the findings are inconsistent. Some studies reported OXT increases after musical activities (; ), while others observed no significant change () or even decreases (; ; ). The discrepancies may arise from differences in participant characteristics, types of music-based intervention, study designs, OXT measurements, and contextual factors (; ; ). Besides, most studies have been conducted in laboratory or clinical settings with limited ecological validity, which raises questions about whether music-induced OXT changes observed under controlled conditions can be generalized to real-life contexts, especially collective musical rituals.

To address the gaps, this systematic review synthesizes current evidence on how music modulates the OXTergic system in humans. By considering the contextual factors, population characteristics, and intervention features, the review refines the theoretical model of music-induced OXT response to offer guidance to optimize music-based interventions for populations with social impairments. For a rapidly developing field, it will provide a reliable foundation for future research and clinical applications.

The objective of the current review is to explore how individual and group musical activities influence mental health through the OXTergic system in terms of social bonding and emotion regulation. We will address the following questions:

  • (1) How do OXT levels change in different populations in music-based interventions?

  • (2) Which characteristics of music-based interventions can be associated with significant changes in OXT levels and psychosocial outcomes?

  • (3) What are the strengths and limitations of OXT measurements in music-based interventions?

2 Methods

2.1 Search strategy

The systematic review was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines. Literature searches were conducted on November 23, 2024, in five databases: PubMed, Embase, Scopus, Web of Science Core Collection, and PsycInfo. Search terms combined keywords related to music, ritual, OXT, emotion regulation, and social bonding. A detailed search strategy is presented in the Supplementary material S1.

2.2 Eligibility criteria

Studies were included if they met the following criteria: (1) original, peer-reviewed, full-text articles published in English up to October 2024; (2) study designs were randomized controlled trials (RCTs), quasi-experimental designs, or observational studies; (3) populations including healthy individuals, clinical populations, or both; (4) interventions involving music therapy or musical activities (e.g., music listening, singing, or playing instruments) delivered individually or in groups; and (5) outcomes reporting OXT measurements along with at least one psychosocial, clinical, or subjective outcome. Exclusion criteria included: (1) non-human studies; (2) reviews or editorials without original data; (3) studies focusing solely on pharmacotherapy (unless compared to music-based interventions); (4) conference abstracts or book chapters; and (5) non-peer-reviewed papers. There were no restrictions regarding participant age, sample size, geographical location, or comparator.

2.3 Study selection

All identified records were imported into reference management software, and duplicates were removed. Titles and abstracts were screened independently by the first author. Full texts of potentially relevant articles were then assessed for eligibility. Uncertainties were resolved through discussion with the second author.

2.4 Quality assessment

The quality of included studies was assessed using the Joanna Briggs Institute (JBI) critical appraisal tools for RCTs and for quasi-experimental studies. The tools evaluate the validity of study design, methodology, statistical conclusions, and risk of bias (, ). The quality assessment was conducted by the first author and verified by the second author.

2.5 Data extraction, analysis, and synthesis

From the full texts of selected articles and their supplementary material, one reviewer extracted the data, including study characteristics, participant characteristics, intervention details, outcome measures, results, and statistical analyses.

Quantitative OXT changes were transformed into qualitative categories: increase, decrease, or no significant change. Psychosocial outcomes were grouped into broader domains: anxiety, depression, empathy, social bonding, trust, stress and relaxation, as well as emotion and mood. Contextual descriptions were extracted and categorized into four dimensions: (1) nature of the activity (e.g., passive listening, active music-making, or music therapy); (2) stress signals (e.g., levels of physical or mental stress); (3) social cues (e.g., presence of others); and (4) familiarity (e.g., familiarity with the music, participants, or environments).

The reviewer conducted a framework synthesis and a narrative synthesis. The framework synthesis was based on the model proposed by , following a five-stage process (): familiarization, framework selection, indexing, charting, and mapping/interpretation. The narrative synthesis was used to summarize patterns within and across studies, explore relationships, and assess the robustness of findings in terms of methodological quality (). Finally, a convergent design () was used to integrate quantitative and qualitative data.

3 Results

3.1 Study selection and study characteristics

A PRISMA flow diagram is presented in Figure 1, showing study identification and inclusion of eligible studies. A total of 1,865 records were identified, of which 1,860 were from electronic databases, and five were from other sources. After removing duplicates, 902 records were screened based on titles and abstracts. Thirty-one full-text articles were assessed for eligibility, and 11 studies were excluded (five duplicates, two conference abstracts, two with wrong outcomes, and two without OXT data). Finally, 20 studies (seven RCTs and 13 quasi-experiments) were included in the systematic review.

Figure 1

Table 1 summarizes the characteristics of the included studies, which were published between 2003 and 2023 across 11 countries. The majority of the studies were conducted in Europe (n = 12, 60%), followed by North America (n = 5, 25%), and Asia (n = 3, 15%). All studies comprised 877 participants. Sample sizes ranged from 4 to 193 participants (median = 25.5).

Table 1

Author (year)Study designCountrySampleIntervention and comparisonMusic genreOutcomeFinding
Quasi-experimentAustria71 healthy adults (26 M, 45 F) aged 17–28 (mean: 23.1 years); Some participants took part in more than one interventionOne session, 20 min: Singing together (n = 37); Singing alone (n = 25); Speaking together (n = 31); Speaking alone (n = 19)Classical choral (“The Armed Man” by Karl Jenkins)Salivary OXT levels; PANAS; IOS scaleOXT decreased less in both group and solo singing compared to group and solo speaking, indicating a stronger role of singing in promoting social bonding.
Quasi-experimentUSA8 ASD children (7 M, 1 F) aged 6–17 (mean: 11.3 years)38 rehearsals from 1 day/week increasing to 3–4 days/week, 2 h per session, for 3 months: SENSE therapyChildren's song (Musical theatrical production of Disney's “The Jungle Book”)Plasma OXT levels; NEPSY-memory for faces; NEPSY-affect recognition; NEPSY-theory of mindSENSE participants showed improvement in face identification and theory of mind skills without significant changes in baseline OXT levels after a long-term intervention.
Quasi-experimentUSA21 adults: 13 WS (6 M, 7 F) aged 19–42 (mean: 30.2 years); 8 TC (4 M, 4 F) aged 19–45 (mean: 29.4 years)One session, 2 h, with stimuli at specific time points: The favorite music that elicited positive emotions, lasted 5–8 min; Cold water at 15 °C, lasted less than 45 sChosen by the participantsSerum OXT levels; Serum AVP levels; Adolph's approachability; Salk Institute Sociability Questionnaire; Scales of Independent Behavior-RevisedCompared with TC group, music and cold stimuli cause an exaggerated release of OXT and AVP in WS group. Higher basal OXT levels were correlated with increased approach to strangers and decreased adaptive social behaviors in WS patients.
Crossover quasi-experimentFinland62 healthy women aged 18–39, divided into high-empathy (n = 32; mean: 24.2 years) and low-empathy (n = 30; mean: 24.9 years) groupsOne session, 60 min: Listening to unfamiliar sad instrumental music, 20 min; Silence, 20 minClassical (Mozart's Piano Concerto No. 23 2nd movements, Kamen's “Discovery of the Camp,” and Piazzola's “Oblivion”)Plasma OXT levels; Profile of mood states; VAS for felt pleasure, being moved, anxiousness, relaxation, and sadnessHigh-empathy participants experienced greater positive mood and feelings of being moved during music exposure compared with the low-empathy participants. In the high-empathy group, OXT levels were significantly lower in the music condition than in the silence condition.
Quasi-experimentUK193 adults (20% M, 80% F): 55 cancer patients (mean age: 60.8 years); 72 current carers (mean age: 56.9 years); 66 bereaved carers (mean age: 59.7 years)One session, 70 min: Choir rehearsal with live musicUnclearSalivary OXT levels; Salivary beta-endorphin levels; Salivary cortisol levels; VAS of mood scales; VAS of stress scales; VAS of connectednessChoir singing reduced salivary OXT, beta-endorphin, and cortisol levels, and increased cytokine activity. Participants with worse mental health experienced greater emotional improvement after singing.
Crossover RCTItaly20 preterm infants (11 M, 9 F) in the NICU: stable medical condition; mean gestational age at test: 34.8 weeks; mean weight at test: 2264 gPreterm infants received 3 conditions during painful procedures, each for 10 min: Maternal singing; Maternal speaking; Standard careUnclearSalivary OXT levels; Premature Infant Pain ProfileFor preterm infants, maternal speech significantly reduced pain scores and increased OXT levels. Maternal singing led to marginally significant increases in OXT levels but did not significantly reduce pain scores.
Crossover RCTItaly20 mothers of preterm infants, aged 23–34 (mean: 29.2 years)2 sessions, each for 10 min: Maternal singing to the baby; Maternal speaking to the babyUnclearSalivary OXT levels; STAI-SMaternal singing and speaking during their infant's painful procedures significantly increased maternal salivary OXT levels and reduced maternal anxiety.
Crossover quasi-experimentCanada8 healthy older adults (1 M, 7 F): cognitively intact; mean age: 72.8 years2 sessions, each for 45 min: Group singing led by a director; Individual singing led by pre-recorded choir activitiesUnclearSalivary OXT levels; PANASGroup singing triggered the OXT release and improved moods, which may result from social factors rather than from individual factors.
Quasi-experimentSweden16 adults: 8 professional singers (4 M, 4 F) aged 26–49 (mean: 36.4 years); 8 amateur singers (2 M, 6 F) aged 28–53 (mean: 40.2 years)One session, 45 min: A classical singing lesson with the participants' teachersClassicalSerum OXT levels; VAS for the dimensions: sad-joyful, anxious-calm, worried-elated, listless-energetic, tense-relaxedBoth amateurs and professionals showed OXT increases after singing. A singing lesson promoted wellbeing in amateurs but not in professionals.
Crossover quasi-experimentSlovak14 healthy males aged 21–292 sessions (washout: 7 days), each for 45 min: Listening to pleasant rock music (played forward) during stress tasks; Listening to unpleasant rock music (played backward) during stress tasksRock (Music excerpts by Rolling Stones)Plasma OXT levels; Plasma ACTH levels; Plasma cortisol levels; Plasma AVP levels; STAI-SThe acute increase in state anxiety induced by unpleasant music modulated physiological responses under stress conditions, with reduced HPA activity and increased OXT and AVP levels, while cortisol levels remained unchanged.
Quasi-experimentUSA4 university students (2 M, 2 F) aged ≥18; Jazz vocalistsOne session, including 2 performances (washout: 30 min): Standard group singing, 5.6 min; Improvised group singing, 6 minJazz (A pre-composed song by the researcher)Plasma OXT levels; Plasma ACTH levels; Flow State Scale-2Both standard and improvised group singing induced social flow in participants. OXT levels decreased in the standard performance, but increased in the improvised condition.
Quasi-experimentGermany21 adults (5 M,16 F) aged 18–65 (Median: over 50 years); Mixed group of novice and experienced chorists, with 9 participants reporting chronic diseases10 sessions, 30 min per session: Choral singing (7th session); Chatting (8th session)Folk rock (four-part choral arrangements of “California Dreamin”)Salivary OXT levels; Ad hoc questionnaire on positive and negative feelingsSinging and chatting significantly increased positive feelings. Salivary OXT levels increased significantly and negative feelings were reduced more after choral singing but not after chatting.
Cluster RCTIndonesia61 healthy fourth-grade elementary students: Playing angklung (12 M, 8 F); Practicing silence (9 M, 12 F); Control group (11 M, 9 F)15 min every day, for 8 weeks, before morning classes: Playing angklung; Practicing silence; Control: no interventionIndonesian traditional music (national and regional songs)Salivary OXT levels; PANAS-Child FormSalivary OXT levels significantly increased in the silence group compared to those in the control group, without significant changes in clinical parameters. A slight decrease in OXT level was observed in the angklung group, possibly due to the cognitive demands of decoding, memorizing, and playing musical notes. The control group showed the greatest OXT decrease.
RCTSweden40 patients who have undergone heart surgery on postoperative day one: Music group (17 M, 3 F; mean age: 64 years); Control group (15 M, 5 F; mean age: 67 years)One session, 60 min: Music group: listening to 30-min soft, relaxing music during bed rest and then 30-min standard bed rest; Control group: 60-min standard bed restAmbient (MusiCure)Serum OXT levels; Relaxation numeric rating scaleBaseline serum OXT levels and relaxation scores were significantly lower in the music group than in the control group, likely due to differences in surgical duration and type, but no differences were observed after intervention. Music during bed rest increased OXT levels, whereas bed rest alone decreased them, suggesting music-related psychological processes may promote OXT release.
Quasi-experimentJapan26 healthy males aged 21–34 (mean: 29.4 years); Without musical training and a habit of listening to classical music2 sessions, each for 20 min: Listening to slow-tempo music sequences; Listening to fast-tempo music sequencesClassical (Piano pieces by Chopin)Salivary OXT levels; Salivary cortisol levels; Heart rate variabilityListening to slow-tempo music significantly increased salivary OXT levels and lnHF, and significantly decreased the heart rate, suggesting an involvement of vagal activity. No correlation was found between OXT changes and self-reported relaxation.
RCTUSA25 adults with post-stroke unilateral hemiparesis: Predominantly Black; MULT-I (5 M, 8 F; mean age: 61.2 years); HEP (8 M, 4 F; mean age: 61.8 years)45 min per session, twice a week, for 6 weeks: Music Upper Limb Therapy-Integrated (MULT-I); Home exercise program (HEP)UnclearSerum OXT levels; Serum BDNF levels; Patient Health Questionnaire-9MULT-I significantly reduced depression levels and increased BDNF levels compared with HEP. There were no significant changes in serum OXT levels in both groups.
RCTAustria30 healthy male graduate and doctoral students: Music group (n = 15; mean age: 27.2 years); Control group: (n = 15; mean age: 26.1 years)One session including 40 rounds of the trust game: Music group: listening to music during the game; Control group: no musicClassical (Schubert's “Marche Militaire”)Plasma OXT levels; Investment amount in the trust game; Trustworthiness 7-point Likert scaleMusic did not have significant effects on OXT levels, trust behavior, or perceived trustworthiness. In the no-music condition, OXT increase was associated with greater perceived trustworthiness but did not influence trust behavior.
Quasi-experimentGermany38 healthy adults: Cohort 1 (9 M, 12 F) aged 19–26 (median: 22 years); Cohort 2 (8 M, 9 F) aged 18–29 (median: 23 years); Student chorists2 sessions, each for 20 min: Cohort 1: A-B (washout: 2 days) or B-A (washout: 4 days); Cohort 2: A-B (washout: 4 days); A: choral singing, B: solo singingClassical (Cohort 1: Handel's “Messiah”; Cohort 2: Bach's “Christmas Oratorio”)Salivary OXT levels; STADI-SBoth choral and solo singing improved moods. Choral singing reduced OXT levels, while solo singing modestly increased OXT levels, which may be caused by increased or decreased stress signals.
Wulff et al. (2021)RCTGermany172 pregnant women aged 18–42 (mean: 34.0 years): Music group (n = 64; gestational age: 31.8 weeks); Singing group (n = 59; gestational age: 30.6 weeks); Control group (n = 49; gestational age: 32.7 weeks)Music: one 30-min session of music listening with a therapist, and home-based music listening for 10–15 min daily until childbirth; Singing: 30-min singing sessions with a therapist 2–4 times, and home-based singing for 10–15 min daily until childbirth; Control: no interventionListening group: Classical; Singing group: Children's songSalivary OXT levels; Self-Assessment Manikin; VAS of perceived closeness; Maternal Antenatal Attachment Scale; The Edinburgh Postnatal Depression ScaleThere were significant increases in salivary OXT levels and reductions in salivary cortisol levels in the singing and music groups. Compared to the music group, the singing group showed greater improvements in emotions, cortisol levels, and perceived closeness to the unborn child. No significant effects were found on depressive symptoms or bonding questionnaire scores.
Yuhi et al. (2017)Quasi-experimentJapan27 maltreated children (22 M, 5 F) aged 8–15Group taiko drumming, less than once a week, for 18 months: Recital: 12 times, 5–60 min (mean: 14.1 min); Practice: 5 times, 80–155 min (mean: 108 min); Free condition: 6 times, 90–200 min (mean: 118 min)Japanese kumi-daikoSalivary OXT levels; Personalities recorded by teachersGroup drumming supported emotional and social improvements in maltreated children, but significant OXT increases were observed only in elementary boys in recital sessions.

Study characteristics.

ACTH, adrenocorticotropic hormone; ASD, autism spectrum disorder; AVP, arginine vasopressin; BDNF, brain-derived neurotrophic factor; F, female; HPA axis, hypothalamic-pituitary-adrenal axis; IOS, Inclusion of Other in the Self; lnHF, logarithmic transformation of low frequency powers of heart rate variability; M, male; NEPSY, A Developmental Neuropsychological Assessment; OXT, oxytocin; PANAS, Positive and Negative Affect Schedule; RCT, randomized controlled trial; SENSE, Social Emotional NeuroScience Endocrinology; STADI-S, the state subscale of State-Trait Anxiety-Depression Inventory; STAI-S, the state subscale of State-Trait Anxiety Inventory; TC, typical control; VAS, visual analog scale; WS, Williams syndrome.

The reviewed music-based interventions encompassed singing, playing instruments, listening to music, and music therapy. Most studies used pieces of classical music as materials. Only two studies adopted quasi-ritualistic music-based interventions grounded in culture: one study from Japan used group taiko drumming (Yuhi et al., 2017), and one study from Indonesia examined a school-based intervention program combining Indonesian traditional music ().

One reviewer used the JBI critical appraisal tools to assess study quality. Among the studies, 17 were rated as low risk of bias, two as moderate risk, and one as high risk. Tables 2, 3 present the detailed results of the quality assessment. No studies were excluded based on quality alone.

Table 2

Author (year)1. Was true randomization used for the assignment of participants to treatment groups?2. Was allocation to treatment groupsconcealed?3. Were treatment groups similar at thebaseline?4. Were participants blind to treatment assignment?5. Were those delivering the treatment blind to treatment assignment?6. Were treatment groups treated identically other than the intervention of interest?7. Were outcome assessors blind to treatment assignment?8. Were outcomes measured in the same way for treatment groups?9. Were outcomes measured in a reliable way?10. Was follow-up complete and if not, were differences between groups in terms of their follow-up adequately described and analyzed?11. Were participants analyzed in the groups to which they were randomized?12. Was appropriate statistical analysisused?13. Was the trial design appropriate and any deviations from the standard RCT design accounted for in the conduct and analysis of the trial?
aYesYesYesN/AN/AYesNoYesYesYesYesYesYes
aYesYesYesN/AN/AYesNoYesYesYesYesYesYes
YesYesYesN/AN/AYesUnclearYesYesYesYesYesYes
YesYesNoN/AN/AYesNoYesYesYesYesYesYes
YesYesNoN/AN/AYesUnclearYesYesYesYesYesYes
YesYesYesN/AN/AYesNoYesYesYesYesYesYes
Wulff et al. (2021)YesYesNoN/AN/AYesNoYesYesYesYesYesYes

Quality assessment for RCTs.

Question 4 and question 5 were marked as N/A because blinding participants and treatment providers was not feasible. For RCTs, studies scoring ≥ 9 on the 13-question checklist were considered high quality, while those scoring < 4 were deemed low quality.

aThe two articles were the same clinical trial: reported data from preterm infants, and focused on the mothers of the preterm infants.

Table 3

Author (year)1. Is it clear in the study what is the cause and what is the effect?2. Was there a control group?3. Were participants included in any comparisons similar?4. Were the participants included in any comparisons receiving similar treatment/care, other than the exposure or intervention of interest?5. Were there multiple measurements of the outcome, both pre- and post-intervention/ exposure?6. Were the outcomes of participants included in any comparisons measured in the same way?7. Were outcomes measured in a reliable way?8. Was follow-up complete and if not, were differences between groups in terms of their follow-up adequately described and analyzed?9. Was appropriate statistical analysis used?
YesNoYesYesYesYesYesYesYes
YesNoYesUnclearYesYesYesYesYes
YesNoNoYesYesYesYesYesYes
YesYesNoYesYesYesYesYesYes
YesNoNoYesYesYesYesYesYes
YesNoYesYesYesYesYesNoYes
YesNoNoNoYesYesUnclearYesYes
YesNoYesYesYesYesYesYesYes
YesNoYesYesYesYesUnclearYesYes
YesNoYesYesYesYesUnclearNoNo
YesNoYesYesYesYesYesYesYes
YesNoYesYesYesYesYesYesYes
Yuhi et al. (2017)YesNoNoNoYesYesNoNoNo

Quality assessment for quasi-experimental designs.

For quasi-experimental studies, those scoring ≥ 7 on the 9-question checklist were rated as high quality, and those scoring < 4 were considered low quality.

3.2 Population characteristics and OXT responses

The included studies encompassed diverse populations, with a primary focus on healthy adults. Other studies examined clinical populations, including cancer patients and their carers , pregnant women Wulff et al., (2021), preterm infants and their mothers , individuals with Williams syndrome , post-surgical patients (), post-stroke patients (), children with autism spectrum disorder (ASD; ), and maltreated children (Yuhi et al., 2017). Two studies (; ) recruited mixed samples that varied in age, sex, health status, and singing experiences, as detailed in Table 1.

Participant ages ranged from 34.8 gestational weeks (preterm infants) to 72.8 years (healthy older adults). Most studies had a predominance of female participants. Baseline peripheral OXT levels showed notable differences between populations. In typical control individuals, the patterns of plasma OXT response to music demonstrated low inter-individual variability (), although they had different baseline OXT levels. By contrast, patients with Williams syndrome exhibited higher baseline plasma OXT levels and greater variability in response to music . In pregnant and postpartum mothers (; Wulff et al., 2021), baseline peripheral OXT levels were low, with small yet statistically significant changes after musical activities. The patterns of OXT response to music were not influenced by sex (; ). Other covariates the phase of the menstrual cycle, the use of hormonal contraception, and the musical sophistication index did not change the patterns ().

Tables 4, 5 summarize the peripheral OXT changes induced by music across populations and contexts (also see Supplementary material S2). These results revealed the context-dependence of OXT responses to music. After group singing, OXT levels decreased in healthy young adults ; ; but increased in healthy older adults . OXT reductions were observed in cancer patients and their carers after choral singing . For preterm infants and their mothers , as well as pregnant women Wulff et al., (2021), maternal singing to infants or fetuses increased their OXT levels. Listening to slow-tempo music generally triggered OXT release in several populations, including healthy males , pregnant women with trait anxiety Wulff et al., (2021), and post-surgical patients . Interestingly, listening to sad music led to significant OXT decreases in healthy females with high empathy but not in healthy females with low empathy . However, in children with ASD and post-stroke patients , no significant changes in peripheral OXT levels were observed after long-term group music therapy.

Table 4

PopulationPeripheral OXT response in different contexts
Group singingImprovised group singingSolo singingListening to slow musicListening to fast musicListening to sad musicPlaying the instruments in a group
Healthy young adults
Healthy females with high empathy
Healthy females with low empathy
Healthy males
↓ or ↑a
Healthy older adults
Healthy children
b

Peripheral OXT changes induced by music in healthy populations.

↑ = significant increase; ↓ = significant decrease; → = no significant change. Significance levels (e.g., p < 0.05, p < 0.01, p < 0.001) are not distinguished in the review. All significance is defined as p < 0.05.

aParticipants showed OXT decreases after listening to pleasant rock music during stress tasks, but showed OXT increases after listening to unpleasant rock music during stress tasks.

bOXT measurements were obtained before and after a long-term program.

Table 5

PopulationPeripheral OXT responses in different contexts
Group singingClassical singing lessonMaternal singing to a babyListening to slow musicListening to the favorite positive musicPlaying the instruments in a groupGroup music therapy
Williams Syndrome (WS)
Typical controls paired with WS
Cancer patients
Carers of cancer patients
Pregnancy with trait anxiety
Wulff et al. (2021)
Mothers of preterm infants
Preterm infants
c
Post-surgical patients
Children with autism spectrum disorder
e
Maltreated boys
Yuhi et al. (2017)d
Post-stroke patients
ee
Mixed sample
b
a

Peripheral OXT changes induced by music in clinical populations.

↑ = significant increase; ↓ = significant decrease; → = no significant change. Significance levels (e.g., p < 0.05, p < 0.01, p < 0.001) are not distinguished in the review. All significance is defined as p < 0.05.

aThe sample included mixed age, sex, clinical conditions, and singing experiences.

bThe sample included mixed age, sex, and singing experiences.

cThe change was marginally significant.

dThere was a high risk of bias due to inconsistency in the duration of intervention and control conditions, the measurement time points, the sample size, and the number of sessions.

eThe measurements were obtained before and after a long-term program.

3.3 Psychosocial outcomes and OXT changes

Across the included studies, associations between psychosocial outcomes and peripheral OXT changes varied by population, context, and measurement time point. Table 6 summarizes the relationships between psychosocial outcomes, OXT changes, and context characteristics.

Table 6

Author (year)PopulationPsychosocial outcomeOXT responseContext
MeasureChangeMeasureChangeActivityStress signalSocial cueFamiliarity
Anxiety
Healthy young adultsSTADI-SSalivary OXT levelsChoral singingPossibleLimitedHigh
Healthy young adultsSTADI-SSalivary OXT levelsSolo singingPossibleNoHigh
Healthy malesSTAI-SPlasma OXT levels (iAUC)Listening to unpleasant rock music during stress tasksHighNoLow
Healthy malesSTAI-SPlasma OXT levels (iAUC)Listening to pleasant rock music during stress tasksHighNoModerate
Mothers of preterm infantsSTAI-SSalivary OXT levelsMaternal singingPossibleYesLow
Depression
Healthy young adultsSTADI-SSalivary OXT levelsChoral singingPossibleLimitedHigh
Healthy young adultsSTADI-SSalivary OXT levelsSolo singingPossibleNoHigh
Post-stroke patientsPHQ-9Plasma OXT levelsGroup music therapy (MULT-I), 6 weeksModerateYesIncreased gradually
Empathy
Healthy females with high empathyVAS-being moved67.2/100aPlasma OXT levelsListening to unfamiliar sad instrumental musicPossibleNoLow
Healthy females with low empathyVAS-being moved52.3/100aPlasma OXT levelsListening to unfamiliar sad instrumental musicPossibleNoLow
ASD childrenNEPSY-affect recognitionPlasma OXT levelsMusical theater therapy, 3 monthsModerateYesIncreased gradually
ASD childrenNEPSY-theory of mindPlasma OXT levelsMusical theater therapy, 3 monthsModerateYesIncreased gradually
Social bonding
Healthy young adultsIOS scaleSalivary OXT levelsChoral singingPossibleYesHigh
Jazz vocalistsFSS-2Presence of social flowaPlasma OXT levelsStandard choral singingModerateYesLow
Jazz vocalistsFSS-2Presence of social flowaPlasma OXT levelsImprovised choral singingHighYesLow
Cancer patientsVAS-connectednessSalivary OXT levelsChoral singingPossibleYesModerate
Carers of cancer patientsVAS-connectednessSalivary OXT levelsChoral singingPossibleYesModerate
ASD childrenNEPSY-memory for facesPlasma OXT levelsMusical theater therapy, 3 monthsModerateYesIncreased gradually
Trust
Healthy malesInvestment amount in the trust gameHigher valuebPlasma OXT levelsListening to fast classical musicPossibleYesLow
Healthy males7-point Likert scale of trustworthinessHigher valuebPlasma OXT levelsListening to fast classical musicPossibleYesLow
Stress and relaxation
Healthy females with high empathyVAS- relaxation79.4/100aPlasma OXT levelsListening to unfamiliar sad instrumental musicPossibleNoLow
Healthy females with low empathyVAS- relaxation62.6/100aPlasma OXT levelsListening to unfamiliar sad instrumental musicPossibleNoLow
Post-surgical patientsRelaxation numeric rating scaleSerum OXT levelsListening to soft, relaxing, slow musicLowNoLow
Mixed sample-amateur singersVAS-tense-relaxedRelaxation ↑Serum OXT levelsClassical singing lessonModerateYesHigh
Mixed sample-professional singersVAS-tense-relaxedRelaxation ↑Serum OXT levelsClassical singing lessonModerateYesHigh
Cancer patientsVAS-stressSalivary OXT levelsChoral singingPossibleYesModerate
Carers of cancer patientsVAS-stressSalivary OXT levelsChoral singingPossibleYesModerate
Emotion and mood
Healthy young adultsPANAS (positive affect)Salivary OXT levelsChoral singingPossibleYesHigh
Healthy young adultsPANAS (positive affect)Salivary OXT levelsSolo singingPossibleNoHigh
Healthy older adultsPANAS (mood rating)Salivary OXT levelsChoral singingUnclearYesHigh
Healthy older adultsPANAS (mood rating)Salivary OXT levelsSolo singingUnclearNoModerate
Healthy childrenPANAS-CSalivary OXT levelsPlaying the angklung, 8 weeksModerateYesIncreased gradually
Healthy females with high empathyPOMS (positive mood)Plasma OXT levelsListening to unfamiliar sad musicPossibleNoLow
Healthy females with low empathyPOMS (positive mood)Plasma OXT levelsListening to unfamiliar sad musicPossibleNoLow
Cancer patientsVAS-moodSalivary OXT levelsChoral singingPossibleYesModerate
Carers of cancer patientsVAS-moodSalivary OXT levelsChoral singingPossibleYesModerate
Wulff et al. (2021)Pregnant mothers with trait anxietySAMValence ↑, Arousal ↓, Dominance ↑Salivary OXT levelsListening to calm classical instrumental musicLowYesUnclear
Wulff et al. (2021)Pregnant mothers with trait anxietySAMValence ↑, Arousal ↓, Dominance ↑Salivary OXT levelsSinging children's songs and lullabiesLowYesModerate
Mixed sample-amateur singersVAS-anxiousness-calmCalm ↑Serum OXT levelsClassical singing lessonModerateYesHigh
Mixed sample-professional singersVAS-anxiousness-calmCalm →Serum OXT levelsClassical singing lessonModerateYesHigh
Participants varied in age, sex, clinical conditions, and singing experiencesAd hoc questionnaire of subjective feelingsPositive feelings ↑, Negative feelings ↓Salivary OXT levelsChoral singingModerateYesVaried in participants

Relationships between psychosocial outcomes, OXT changes, and context characteristics.

↑ = significant increase; ↓ = significant decrease; → = no significant change. Significance levels (e.g., p < 0.05, p < 0.01, p < 0.001) are not distinguished in the review. All significance is defined as p < 0.05.

ASD, autism spectrum disorder; FSS-2, Flow State Scale-2; iAUC, incremental Area Under the Curve; IOS, Inclusion of Other in the Self; MULT-I, Music Upper Limb Therapy-Integrated; NEPSY, A Developmental Neuropsychological Assessment; OXT, oxytocin; PANAS, Positive and Negative Affect Schedule; PANAS-C, Positive and Negative Affect Schedule-Child Form; PHQ-9, Patient Health Questionnaire-9; POMS, Profile of Mood States questionnaire; SAM, Self-Assessment Manikin; STADI-S, the state subscale of State-Trait Anxiety-Depression Inventory; STAI-S, the state subscale of State-Trait Anxiety Inventory; VAS, visual analog scale.

aThe outcome was measured only once after the intervention.

bThe outcome showed a higher value in the music condition than in the no-music condition, but the difference was not statistically significant.

Long-term intervention programs did not lead to significant changes in peripheral OXT levels, despite the improvements in psychosocial outcomes. For instance, after participating in interactive improvised group music-making therapy for 6 weeks, post-stroke patients showed improvements in depression and motor functions without significant OXT changes . Similarly, a 3-month structured musical theater therapy improved the ASD children's skills of theory of mind and memory for faces, but no significant plasma OXT changes were observed .

Short-term music-based interventions, such as a single session of singing, often resulted in improvements in anxiety, though not always accompanied by changes in peripheral OXT levels. Maternal singing led to anxiety reduction and OXT increases in mothers of preterm infants , whereas in other contexts (e.g., healthy adults engaging in choir singing), anxiety reductions occurred accompanied by OXT decreases . Additionally, unpleasant music stimuli under stress conditions both increased anxiety and plasma OXT levels in healthy males .

Empathy-related outcomes unveiled noteworthy patterns. For high-empathy healthy females, listening to unfamiliar sad instrumental music elicited stronger self-reported emotional responses, accompanied by significant decreases in plasma OXT levels. In contrast, low-empathy healthy females reported smaller emotional improvements but showed modest OXT increases .

In an RCT, examined whether music listening in a trust game influenced the trust behavior and perceived trustworthiness. During the trust game, no significant changes in OXT levels were observed in the participants in the music conditions, but they tended to show greater trust behavior and rated faces and nicknames as more trustworthy compared to those in the no-music conditions, although the trends did not reach statistical significance.

All studies on choral singing reported that participants had experienced social bonding, and OXT levels declined in most cases ; ; , except in improvised choral singing, where OXT levels increased .

Last, participants frequently reported subjective improvements in relaxation and mood scales after musical activities across studies. However, these changes were not always accompanied by increases in peripheral OXT levels ; ; ; ; ; ; ; ; Wulff et al., (2021).

Overall, improvements in psychosocial outcomes do not correspond with changes in peripheral OXT levels, and consistent OXT response patterns were observed only for specific psychosocial outcomes.

3.4 Effects of contextual factors on patterns of OXT responses and psychosocial outcomes

3.4.1 Patterns of OXT responses

Analyses of OXT responses across studies consistently revealed interaction effects, which indicated that patterns of OXT change were context-dependent (see Table 7). demonstrated a significant time × condition interaction for salivary OXT responses (F(1,21) = 7.988, p < 0.05). The dependence on activity was further supported by several studies: reported a significant time × condition interaction (χ2(2) = 6.99, p = 0.03), as did (z = −2.142, p = 0.032). observed a significant time × context interaction (F(4) = 7.27, p < 0.001), and found a significant time × vocal mode interaction (χ2(1) = 5.7, p = 0.018). Evidence for OXT's sensitivity to stress signals was provided by , who observed a significant time × tempo interaction (F(1,22) = 13.44, p = 0.0014). However, found no significant time × social context interaction (p = 0.230) or time × vocal mode × social context interaction (p = 0.190), indicating that OXT responses were not significantly influenced by the explicitly manipulated social setting in their study.

Table 7

Author (year)OutcomeStatistical analysisContextual factor
InteractionResult
Physiological outcome
Salivary OXT levelsTime × vocal modeχ2(1) = 5.7, p = 0.018Activity (singing vs. speaking)
Salivary OXT levelsTime × social contextp = 0.230 (non-significant)Social cue (together vs. alone)
Salivary OXT levelsTime × vocal mode × social contextp = 0.190 (non-significant)Activity (singing vs. speaking); Social cue (together vs. alone)
Salivary OXT levelsTime × conditionχ2(2) = 6.99, p = 0.03Activity (maternal singing vs. maternal speaking vs. standard care)
Salivary OXT levelsTime × conditionz = −2.142, p = 0.032Activity (choral singing vs. solo singing)
Salivary OXT levelsTime × conditionF(1,21) = 7.988, p < 0.05Activity (choral singing vs. chatting)
Salivary OXT levelsTime × tempoF(1,22) = 13.44, p = 0.0014Stress signal (slow music vs. fast music)
Salivary OXT levelsTime × contextF(4) = 7.27, p < 0.001Activity (choral singing vs. solo singing)
Psychosocial outcome
IOS scaleTime × vocal modeχ2(1) = 7.06, p = 0.008Activity (singing together vs. speaking together)
PANASTime × vocal mode × social contextχ2(1) = 10.54, p = 0.001Activity (singing vs. speaking); Social cue (together vs. alone)
PANASTime × conditionz = −2.321, p = 0.020Activity (choral singing vs. solo singing)
STAI-STime × conditionF(1, 26) = 5.10, p < 0.05Stress signal (unpleasant music vs. pleasant music)
Ad hoc questionnaire of subjective feelings-positive feelingsTime × conditionF(1,20) = 9.655, p < 0.01Activity (choral singing vs. chatting)
STADI-S-excitementTime × contextF(1) = 6.51, p = 0.015Activity (choral singing vs. solo singing)
STADI-S-happinessTime × contextF(1) = 5.27, p = 0.028Activity (choral singing vs. solo singing)

Effects of contextual factors on OXT response patterns and psychosocial outcomes.

This table presents studies that detected significant statistical interactions, which indicates that the temporal patterns of OXT levels or psychosocial outcomes differed across contexts. The “Result” column shows test statistics and p-values for the interaction. All significance is defined as p < 0.05.

IOS, Inclusion of Other in the Self; OXT, oxytocin; PANAS, Positive and Negative Affect Schedule; STADI-S, the state subscale of State-Trait Anxiety-Depression Inventory; STAI-S, the state subscale of State-Trait Anxiety Inventory.

3.4.2 Patterns of psychosocial outcomes

Psychosocial outcomes exhibited varying degrees of context-dependence (see Table 7). found a significant time × condition interaction for positive feelings (F(1,20) = 9.655, p < 0.01). observed significant time × context interactions for excitement (F(1) = 6.51, p = 0.015) and happiness (F(1) = 5.27, p = 0.028), but not for worry or sadness. examined the effects of pleasant vs. unpleasant music and found a significant time × condition interaction for state anxiety (F(1, 26) = 5.10, p < 0.05). In , the social bonding measure (IOS scale) showed a significant time × vocal mode interaction (χ2(1) = 7.06, p = 0.008). In contrast, the subjective emotion measure (PANAS) in the same study exhibited a significant three-way interaction among time, vocal mode, and social context (χ2(1) = 10.54, p = 0.001). The findings indicate that changes in psychosocial outcomes, particularly positive affect measures, were significantly influenced by activity and social cues.

3.5 The characteristics of OXT measurements

All the included studies used peripheral OXT measurements as biomarkers, and they varied in collection protocols, sample types, and quantification methods, as summarized in Table 8. Twelve studies provided details on the time of sample collection. Of these, nine studies scheduled collection during the afternoon or evening hours to avoid interference from the cortisol awakening response (CAR). The remaining studies adopted different protocols: one study collected samples across many time points throughout the day (morning, afternoon, and evening), and another between 9:00 and 13:00. Notably, was the sole study that conducted sample collection during the CAR period in the early morning.

Table 8

Author (year)Intervention durationSample typeCollection timeMeasurement time pointExtractionQuantificationOXT level (pg/ml)
20 minSaliva18:00–20:00; 10 min/samplet = −5 (−10–0), 30 (25–35) minYesEIA45.9–305.5a
120 min (3-month program)PlasmaUnclearBefore and after the entire programUnclearEIAN/A
Music: 5–8 min, Cold water: < 45 sSerum13:00–16:00Baseline: t = −5, 0 min, Experiment: t = 1, 5, 10, 15, 20, 25, 30, 45 minNoEIAWS: 130–7500; TC: 80–300
20 minPlasmaUnclearBaseline: t = 0 min, Experiment 1: t = 20 min, Experiment 2: t = 40 minYesEIA0.59–25.58
70 minSaliva19:00–20:15t = 0, 75 minNoFluorescence bead-based multiplex immunoassay3.62–5.15
10 minSalivaUncleart = 0, 10 minNot requiredRIA0.78–1.4
10 minSalivaUncleart = 0, 10 minNot requiredRIA2.6–3.0
45 minSalivaAt the same time of day (evening)t = 0, 45 minNoElectrochemiluminescence assay7.8–14.8
45 minSerum9:00–20:15Before and 30 min after the lessonNoEIA469–788
45 minPlasma12:15–16:00t = 0, 15, 30, 45, 60, 90 minNot requiredRIAN/A
Nearly 6 minPlasmaUncleart = −5, 6 minNoEIA90–370
30 minSalivaUncleart = 0, 30 minUnclearEIA13.04–18.08
15 min (8-week program)SalivaIn the morningBefore and after the entire programNoEIA82.3–173b
30 minSerum12:00–13:00t = 0, 30, 60 minYesEIA62.6–73.9
20 minSaliva14:00–18:00; 1–3 min/samplet = 0, 20 minYesEIA2.26–17.02
45 min (6-week program)SerumUnclearBefore and after the entire programNoEIA59.26–68.73
During 40 rounds of the trust gamePlasma9:00–13:00Pre-game, post 10 rounds, post 20 rounds, post 30 rounds, post 40 rounds, 10 min after the gameYesFluorescence immunoassay27–51
20 minSaliva18:00–20:30; 1 min/sampleBaseline: t = −20, 0 min, Experiment: t = 10, 20 min, Post-session: t = 40 minNot requiredRIA0–13
Wulff et al. (2021)30 minSaliva13:00–16:00t = 0, 30 minNot requiredRIA1.00–1.11
Yuhi et al. (2017)Recital: 5–60 min, Practice: 80–155 min, Free condition: 90–200 minSaliva2–4 min/sample10 min before and 10 min after each sessionNoEIA83–265a

The characteristics of OXT measurements.

For OXT levels, the values in the table indicate the range described in the references or estimated from charts or their supplementary data. The starting time points of interventions were corrected to 0 min.

EIA, enzyme immunoassay; RIA, radioimmunoassay; TC, typical control; WS, Williams syndrome.

aHigh concentrations are likely due to prolonged collection.

bThe quantitative unit was not reported.

In short-term interventions, samples were often collected immediately before and after a single session to measure acute OXT changes. Six studies collected samples at three or more time points (; ; ; ; ; ), whereas the others relied on pre-test and post-test measurements. Long-term interventions assessed baseline OXT levels before and after the entire program, which lasted from several weeks to months. However, long-term interventions generally showed no significant changes in baseline OXT levels.

Sample types included plasma (n = 5), serum (n = 4), and saliva (n = 11). Generally, plasma and serum samples showed higher OXT concentrations (ranging from tens to hundreds of pg/ml) compared to saliva samples (approximately 0–20 pg/ml). In terms of quantification methods, 12 studies used enzyme immunoassay (EIA), of which four applied sample extraction procedures prior to EIA. Five studies used radioimmunoassay (RIA), one study used electrochemiluminescence assay, and two used fluorescence immunoassay. Studies that omitted sample extraction before EIA tended to report higher OXT concentrations (; ; ; ; ; Yuhi et al., 2017) than those that included extraction steps (; ; ; ).

Therefore, methodological differences in OXT measurements may contribute to inconsistencies in OXT levels reported across studies.

4 Discussion

4.1 The bidirectional modulation of OXT

4.1.1 Context-dependence of OXT responses

High-quality studies have shown that music can lead to both increases and decreases in peripheral OXT levels. Overall, clinical condition appears to be the most influential factor in modulating OXT release. For instance, individuals with Williams syndrome exhibit exaggerated OXT release in response to their favorite positive music, suggesting dysregulation of the OXTergic system (). In pregnant and postpartum mothers, the OXTergic system is highly active yet peripheral OXT levels tend to be low (; Wulff et al., 2021), since pregnancy inhibits OXT secretion but enhances OXT receptor (OXTR) expression and enzymatic degradation (Uvnas-Moberg, 2024).

Context-dependence can explain the variability in OXT responses across studies. Under basal conditions, central and peripheral OXT levels show no significant correlation (; Valstad et al., 2017), since peripheral OXT reflects only the activity of magnocellular OXT neurons projecting to the posterior pituitary but central release involves both magnocellular and parvocellular OXT neurons (). However, during acute stress or strong physiological or social stimulation, coordinated OXT release in the CNS and periphery may occur and produce higher correlations between central and peripheral OXT levels (; Valstad et al., 2017). As illustrated in Figure 2, music can induce such release by shaping the context, which is comprised of several elements, including the type of activity, stress signals, social cues, and familiarity with the music or social setting. The auditory signals travel through the classical pathway to the auditory cortex, while contextual inputs, including all sensory inputs, stress signals, and social cues, go through the non-classical pathway to the limbic system and the hypothalamus. The inputs are integrated in the hypothalamus, which may trigger coordinated or independent release of OXT in the CNS and periphery, and activate or deactivate the hypothalamic-pituitary-adrenal (HPA) axis.

Figure 2

Nevertheless, the dimensions that constitute the context differ in their importance. Our contextual analysis (Tables 6, 7) indicates that activity and stress signals have combined effects on OXT responses, with qualitative analysis showing that activity is the primary driver, while stress signals may reverse these effects under certain conditions. Social cues and familiarity showed limited predictive power. Statistical evidence also supports the patterns. found a significant time × context interaction, showing that OXT changes over time depend on contexts, consistent with other studies (; ; ). distinguished two contextual factors, vocal mode and social context, and identified vocal mode (corresponding to activity type) as the key factor that influences OXT response patterns, with social context showing no significant effect. confirmed OXT's responsiveness to stress signals, as evidenced by a significant time × tempo interaction effect in a music-listening task that manipulated arousal or stress levels via tempo.

As for the weak effect of the social context, we attribute this to potential issues in the current experimental paradigm. Just as HPA axis activation requires high-stress conditions, OXT modulation may need sufficiently demanding social cues. observed the influence of social cues through improvised group singing, which involved trust, synchrony, and real-time interaction. Instead, the structured singing tasks in and required participants to focus on vocal performance with limited spontaneous social interactions and should have been insufficient to reveal the effects of social cues. As a result, the context-dependence of OXT responses to music is specifically attributed to its reliance on two factors, namely the activity type and the stress signals, based on the statistical findings.

4.1.2 Functional adaptation of OXT to musical contexts

OXT exhibits temporally and spatially specific release patterns, and peripheral release can occur independently of or in parallel with central release (). In the circulation, OXT follows a pulsatile release pattern, and pulse characteristics are strongly linked with baseline socio-emotional functioning. Mean OXT pulse height and pulse mass were negatively correlated with avoidant attachment and positively correlated with perceived social support (). It suggested that temporal dynamics of peripheral OXT release carry crucial psychosocial information.

In the CNS, spatial specificity of OXT release allows for nuanced modulation. OXT activates OXTRs, which are coupled with Gq or Gi/o proteins, resulting in diverse downstream effects (). For example, OXT release directly activates dopaminergic neurons and indirectly inhibits them via local GABAergic interneurons, but the relative magnitudes of the two mechanisms differ in the ventral tegmental area (VTA) and substantia nigra pars compacta (SNc; Xiao et al., 2017). OXT functions as a allosteric modulator of μ opioid receptor (MOR), and enhances the MOR signaling without altering receptor affinity (). In the striatum, OXTR activation on astrocytes inhibits glutamate release (), and D2R-OXTR heterocomplexes facilitate dopaminergic signaling (; ). The differential modulation of OXT supports its excitatory and inhibitory effects, which depend on cell type, receptor coupling, and local neurotransmitter interactions.

Thus, OXT increases or decreases play different roles in reward processing, social cognition, and social behaviors:

  • (1) Reward processing: the ventral striatum is central in reward processing. The nucleus accumbens (NAcc) is responsible for recognizing the salience and valence of stimuli (). The left NAcc recognizes pleasure, while the right NAcc responds to all noteworthy stimuli whether they are positive or negative. When OXT increases, it may promote dopaminergic signaling (Xiao et al., 2017) and enhance the output of the striatum.

  • (2) Empathy network: the ventromedial prefrontal cortex (vmPFC), a core area of the empathy network (), normally receives inhibitory OXT projections. Increased OXT reduces activity in vmPFC, accelerates the response to social cues, reduces recall accuracy, and blurs the boundaries between self and others (Zhao et al., 2016). Decreased OXT might enhance the local activity of the empathy network and right striatum, strengthen emotional responses, and promote deeper emotional engagement with music, though direct evidence for this pattern in musical contexts remains to be established.

  • (3) Social behavior: in both approach and avoidance behaviors, increased OXT reduces right striatum activity, potentially suppressing excessive emotional interference and promoting appropriate behavioral selection. Although OXT may weaken overall motivation, it improves the accuracy of approach to positive stimuli (Yao et al., 2018). This effect may reflect a compensatory mechanism in which motivation control shifts toward the left striatum when the right striatal activity is reduced.

The bidirectional modulation of OXT suggests that the OXTergic system adapts flexibly to musical context. Happy music primarily activates the left NAcc (), and OXT release may facilitate dopaminergic signaling (). Sad music selectively activates the right NAcc over the left (), and is often accompanied by peripheral OXT decreases, especially in females with high trait empathy ().

Furthermore, there is a bidirectional relationship between the OXTergic system and the HPA axis (; ; ). Under acute stress conditions, three temporal patterns may occur:

  • (1) Concurrent release of OXT and adrenocorticotropic hormone (ACTH)/cortisol ();

  • (2) An initial OXT increase that attenuates HPA activation (; );

  • (3) HPA activation that selectively inhibits OXT release from the PVN (; ).

Engaging in low-exertion social activities under low-stress conditions is conducive to OXT release since OXT is primarily involved in physiological relaxation rather than psychological relaxation (). Slow-tempo music is more likely to induce OXT release without inhibition by the HPA axis (; Wulff et al., 2021), as it is usually perceived as more pleasant and relaxing (). In contrast, fast-tempo music or singing may initially activate the HPA axis due to increased arousal or stress, which can temporarily suppress OXT release, even though the HPA responses may decrease by the end of the intervention (; ; ; ; ; ; ; ). Additionally, prolonged singing sessions (e.g., over an hour) may lead to fatigue and limit other social interactions (). If the reward from music is insufficient to counteract the earlier inhibition, OXT levels may remain unchanged or even decrease, as observed in studies where the participants sang touching sublime religious songs (; ).

In summary, the findings challenge the assumption that higher OXT levels are beneficial. Instead, both increases and decreases in OXT levels may support adaptive emotional and social functions. Importantly, the context-dependence of music and OXT emphasizes the need to consider baseline physiological and psychological states, intervention types, and environmental factors.

4.2 Peripheral OXT and subjective outcomes

Discrepancies between peripheral OXT levels and psychosocial outcomes have led to ongoing debate about the relationship between central and peripheral OXT activity. While context-dependence may account for the variability and direction of peripheral OXT responses, subjective outcomes exhibit different context-dependence.

Evidence from implied the dissociation. The patterns of OXT responses were mainly determined by vocal mode (time × vocal mode interaction) rather than social cues. IOS scale also showed a significant time × vocal mode interaction, which suggested that social bonding might be associated with the physiological processes behind vocal synchrony. However, PANAS showed a significant three-way interaction (time × vocal mode × social context), indicating that the changes of affective experiences were dependent on both the activity and social cues. In addition, there were no significant correlations between the pre-post change of OXT and self-reported emotion, mood or stress scores (; ). Thus, different outcome measures are sensitive to different contextual factors. OXT responses should be regarded as intrinsic physiological adaptations to musical activities. Since OXT secretory dynamics (e.g., pulse characteristics) are more likely to reflect subjective socio-emotional functioning than a single measurement or pre-post change (), pre-post changes of OXT often fail to predict subjective outcomes, which are heavily influenced by placebo effects, expectations, and the infeasibility of blinding in music-based interventions. As a result, perceived benefits may reflect cognitive and cultural interpretations of musical meaning.

Behavioral changes following musical activities are likely mediated by central mechanisms not fully captured by peripheral OXT measurements (). For example, social cognition and reward processing can be mediated by OXT and OXTR, while social communication relies more on arginine-vasopressin (AVP) and its receptor (V1aR; ). Since OXT and AVP share structural similarities and can activate each other's receptors, this phenomenon, “cross-talk” (), is likely to participate in shaping social behaviors in musical contexts. Therefore, peripheral OXT levels have limited predictive value for subjective outcomes. When interpreting outcomes regarding social cognition, social reward, or social behavior, we recommend considering the OXT-AVP systems that mediate the adaptive responses to social stimuli.

4.3 Clinical implications

The bidirectional modulation of OXT suggests that therapeutic strategies need to move beyond merely increasing OXT levels, as OXT enhances the salience of both positive and negative social cues, encouraging social approach in safe contexts but inducing avoidance in aversive environments (). From a translational medicine perspective, the context-dependence of OXT responses implies that existing music-based interventions could be refined by considering the activity type and stress signals involved.

For social anxiety disorder (SAD) patients, social cues belong to intense stress signals since they usually interpret social cues negatively and lack motivation for positive social interaction. Based on our contextual analysis, at the initial stage, treatment should begin with a safe setting using neutral, calming, or peaceful music, with limited social interactions. It is advisable that the musical emotions not completely simulate anxiety to avoid overwhelming negative feelings in patients. Social engagement can be gradually introduced only after basic trust is established and the emotional state is stabilized. We recommend structured group singing in a small choir () for SAD patients, as it provides appropriately limited social cues which is sufficient for therapeutic effects but not overwhelming.

Some typical characteristics of ASD, such as deficits in imitation, emotional empathy, and attributing intentions to others, reflect the mirror system dysfunction (). For ASD children, rhythmic musical activities could mobilize the mirror systems, improve self-control, and help process acoustic prosody for conversational quality (Wynn et al., 2022). Because rhythmic music involves precise timing prediction and synchronization, patients need to pay more attention to social cues. Patients who are sensitive to OXT will benefit from improvement in social information processing. Thus, group drumming (Yuhi et al., 2017) and musical theater therapy () are recommended for individuals with ASD. The therapy design requires setting up lessons of different levels and breaking them down into manageable parts to learn. Rhythm complexity needs to increase gradually, from simple, accessible patterns to more complex synchronization tasks, to avoid frustrations for beginners.

Understanding the mechanisms through which music influences emotion regulation and social bonding benefits the design of music therapy, as it enables a better understanding of which strategies are more effective in mobilizing the neural networks to achieve therapeutic goals.

4.4 Strengths and limitations of the included study

Several included studies demonstrated methodological rigor that strengthens confidence in their findings. The RCTs provided stronger evidence for causal relationships between music-based interventions and OXT changes. Some studies accounted for biological and environmental confounders, used validated OXT assays with extraction procedures (), and employed real-world settings that enhanced ecological validity. An advantage was the predominant use of saliva samples. Unlike plasma OXT, which fluctuates rapidly due to its 7-min half-life () and pulsatile release pattern (), salivary OXT can provide more stable measurements since it has a longer half-life ().

However, the studies exhibit several limitations that necessitate discussion. A primary concern is the reliance on pre-post measurements of peripheral OXT, which may not reflect central OXT activity (; Valstad et al., 2017) or be correlated with attachment style, perceived social support, and emotional awareness (). Timing issues were also evident in some studies. Specifically, OXT's dramatic fluctuations within CAR period could potentially mask intervention effects. Comparability across studies was hindered by inconsistencies in sampling methods and quantification techniques (), as well as insufficient reporting of musical and contextual details. Moreover, common issues like small sample sizes and the lack of negative controls reduced the internal validity. While fasting and exercise restrictions did not have significant influence on OXT levels (), other behavioral confounders, such as hospital anxiety or sexual activity prior to sampling, still need to be considered ().

4.5 Strengths and limitations of the current study

This review addressed the research questions through comprehensive search strategies, predefined eligibility criteria, and structured syntheses that integrated heterogeneous findings. Our contextual analysis provides a key insight by identifying four distinct contextual dimensions and their different contribution on OXT modulation.

However, there were several limitations in the review. First, it was mainly conducted by one reviewer, which may introduce selection bias, and restrictions to English publications could have excluded relevant studies. None of the included studies examined OXTR gene polymorphisms, which associate with social behavior variability in clinical populations () and may influence individual responses to music. Additionally, no studies directly investigated receptor interactions such as functional complexes of OXTR with dopamine, serotonin, and opioid receptors. Only reported correlated declines of OXT and beta-endorphin after choral singing. Thus, the interpretation scope of the review is limited.

5 Conclusion

The context-dependence of OXT responses to music has significant implications for clinical practice and research. Evidence supports a bidirectional relationship in which music induces subtle emotional responses, which are further modulated by OXT. The effectiveness of music-based interventions may not rely on the direction of OXT change but on the adaptive modulation of OXT. For clinical applications, music-based interventions should be carefully tailored to specific populations and contexts. Larger sample sizes, standardized outcome measures, and well-controlled experimental designs are needed in the future. Given the limitations of peripheral OXT measurements, we suggest that future studies explore central OXT signaling or receptor interactions. Integrating genetic studies and peptide assessments, the use of OXTR agonists or antagonists, combined with neuroimaging techniques such as functional magnetic resonance imaging (fMRI), may provide deeper insights into how music modulates brain activity through the OXTergic system.

Statements

Data availability statement

The original contributions presented in the study are included in the article/Supplementary material, further inquiries can be directed to the corresponding author.

Author contributions

T-HC: Conceptualization, Data curation, Formal analysis, Visualization, Writing – original draft. C-GT: Conceptualization, Supervision, Writing – review & editing.

Funding

The author(s) declared that financial support was received for this work and/or its publication. This work was supported by a grant project [NSC 112-2410-H-002-071] from Ministry of Science and Technology, Taiwan.

Acknowledgments

The article includes content from T-HC's Master's thesis (), submitted to National Taiwan University. The thesis has been deposited in the university repository but is not publicly accessible currently. The present manuscript contains updated analyses and revised discussion compared to the thesis. All final decisions and revisions were made by the authors.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that generative AI was used in the creation of this manuscript. The author(s) used ChatGPT (GPT-4o model, OpenAI) to improve the clarity and readability of the English in the manuscript. The use of this AI tool was limited to language editing and did not affect the scientific content, interpretations, or conclusions.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fcogn.2025.1678665/full#supplementary-material

Supplementary material S1

Systematic review protocol.

Supplementary material S2

Peripheral OXT changes induced by music over time.

References

Summary

Keywords

music therapy, oxytocin, context-dependence, social cognition, affect regulation, social behavior, ritual music

Citation

Chu T-H and Tsai C-G (2026) Music's context-dependent influence on oxytocin, social bonding, and emotion regulation: a systematic review. Front. Cognit. 4:1678665. doi: 10.3389/fcogn.2025.1678665

Received

03 August 2025

Revised

24 November 2025

Accepted

28 November 2025

Published

02 January 2026

Volume

4 - 2025

Edited by

Kumiko Toyoshima, Osaka Shoin Women's University, Japan

Reviewed by

Alan Harvey, University of Western Australia, Australia

Madhavi Rangaswamy, Christ University, India

Updates

Copyright

*Correspondence: Chen-Gia Tsai,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics