ORIGINAL RESEARCH article

Front. Educ., 25 March 2026

Sec. Language, Culture and Diversity

Volume 11 - 2026 | https://doi.org/10.3389/feduc.2026.1788875

Integrating music, gamification, and acoustic visualization enhances prosody and comprehensibility in secondary EFL classrooms

  • 1. Department of General and Specific Didactics, Faculty of Education, University of Alicante, San Vicente del Raspeig, Spain

  • 2. Department of Pedagogy and Didactics, Faculty of Education, University of Santiago de Compostela, Santiago de Compostela, Spain

Abstract

Introduction:

This study investigates how a multimodal pedagogical intervention integrating music, gamified technology, and acoustic visualization influences the development of prosody, speech comprehensibility, and affective variables in English as a foreign language at the lower secondary level.

Methods:

Adopting a mixed-methods quasi-experimental design, the intervention was implemented during regular classroom instruction with intact groups. Quantitative data included acoustic measurements, native-speaker ratings of comprehensibility, and questionnaires assessing motivation, self-efficacy, and pronunciation-related anxiety, collected at three time points: pretest, immediate post-test, and delayed post-test. Qualitative data were gathered through group interviews and classroom observation.

Results:

Results indicate that learners exposed to the intervention achieved significantly greater gains than the control group in prosodic features related to intonation, rhythm, and stress, as well as in overall speech comprehensibility, while no differential improvement was observed in segmental accuracy. Acoustic analyses revealed changes in rhythmic and melodic patterns consistent with a shift toward stress-timed speech, and perceptual evaluations confirmed a substantial increase in ease of understanding for native listeners. In addition, the intervention was associated with increased motivation and communicative self-efficacy, together with reduced pronunciation-related anxiety.

Discussion:

Follow-up data show that a substantial proportion of these gains was maintained three months after the end of the intervention.

1 Introduction

The teaching of English as a foreign language, particularly within secondary school contexts, is currently undergoing a process of paradigmatic redefinition. The traditional approach, largely centered on segmental competence and grammatical accuracy, has been progressively challenged by contemporary perspectives in applied linguistics. These approaches foreground intelligibility and comprehensibility as the core dimensions of effective oral communication, while distinguishing them from accentedness as a related but conceptually separate perceptual construct, and as priority objectives in the teaching of pronunciation within formal educational settings (; ).

Following , it is essential to distinguish between three related but distinct constructs in L2 pronunciation research: intelligibility, defined as the extent to which an utterance is actually understood by a listener; comprehensibility, defined as the listener's perceived ease or difficulty in understanding speech; and accentedness, referring to the perceived deviation from a local or native-like norm. While these constructs are interrelated, they are not interchangeable. The present study focuses specifically on comprehensibility, as suprasegmental features have consistently been shown to exert a more robust influence on perceived communicative ease than on the degree of foreign accent per se.

Within this framework, prosody is no longer regarded as a peripheral component of speech but rather as a decisive factor in communicative effectiveness. Suprasegmental features, including intonation, rhythm, stress, and linking phenomena, play a central role in organizing discourse, structuring units of meaning, and facilitating cognitive processing on the part of the listener (). Research has consistently demonstrated that prosodic competence constitutes a more stable predictor of comprehensibility than accuracy in the articulation of isolated phonemes, particularly in authentic communicative interactions beyond controlled experimental settings (; ).

Despite this broad scientific consensus, systematic instruction in prosody continues to occupy a marginal position within the official curricula of compulsory secondary education. This situation is especially problematic in Spanish-speaking contexts, given the typological distance between the languages involved. Whereas Spanish is commonly described as a syllable-timed language, English displays a stress-timed pattern characterized by hierarchical alternation between stressed and unstressed syllables, together with substantial processes of vowel reduction (). In the absence of targeted instruction and exposure to varied oral input, learners tend to transfer rhythmic patterns from their first language to the foreign language. This transfer often results in monotonous productions, limited melodic variation, and inadequate discourse segmentation, all of which undermine comprehensibility and contribute to negative perceptions of foreign-accented speech (; ).

The limited presence of prosody within compulsory education stands in sharp contrast to its functional relevance in real language use and reflects a persistent gap between pronunciation research and classroom practice. In secondary education, curricular constraints, limited specialist training among teachers, and the lack of accessible pedagogical resources frequently lead to pronunciation being addressed sporadically and reactively, rather than through systematic planning aimed at fostering sustained prosodic awareness. As a result, many students complete compulsory schooling with substantial lexical and grammatical knowledge, yet continue to experience enduring difficulties in producing oral discourse that is prosodically comprehensible and fluent.

This challenge is further intensified when the affective and neuropsychological dimensions of learning during adolescence are taken into account. This developmental stage is characterized by heightened sensitivity to social evaluation and a persistent concern with public self-image, factors that frequently manifest as communicative anxiety and reduced oral participation. Within this context, the so-called affective filter () is particularly readily activated when pronunciation instruction relies on mechanical repetition or on corrective practices perceived as punitive. Low pronunciation-specific self-efficacy consequently fosters avoidance behaviors that restrict the practice required for the automatization of new phonological and prosodic patterns. This dynamic underscores the need for intervention models that address both the technical dimension of prosody and the emotional climate of the classroom through resources perceived as meaningful and motivating ().

Against this backdrop, the convergence of music and language offers a compelling pedagogical avenue, as neurocognitive research has demonstrated that musical and linguistic processing draw on shared neural resources and mechanisms of temporal prediction, particularly in relation to pitch perception and the hierarchical organization of rhythm (; ). Patel's OPERA hypothesis (2011) posits that musical training places greater demands on auditory precision than everyday speech, thereby sharpening the auditory system's capacity to process prosodic nuances in language. Accordingly, the use of popular songs in the classroom can provide rhythmically structured and affectively engaging speech models, while also supporting the consolidation of procedural memory traces through implicit repetition and the embodied experience of rhythm that characterizes such learning approaches (; ; ).

Nevertheless, the use of music in isolation does not in itself guarantee deliberate attention to linguistic form. For this reason, the integration of gamification and acoustic visualization technologies represents a qualitative advancement in the teaching of prosody.

Gamification, through platforms such as LyricsTraining, introduces challenge-based dynamics and immediate feedback that transform phonetic practice into a sustained and enjoyable experience, reduce anxiety associated with oral production, and promote task persistence (; ). Meta-analytic and systematic evidence on gamification in education (, ) highlights the role of structured game-based dynamics in fostering sustained engagement and persistence in learning tasks. Within EFL prosody instruction, this sustained engagement is especially pertinent, given that the internalization of stress-timed rhythmic patterns depends on repeated and attentive practice.

Complementarily, the acoustic analysis software Praat provides structured metacognitive support by enabling learners to visualize the fundamental frequency contour of their own voices. This representation renders pronunciation observable and manipulable, facilitates comparison with reference models and fosters assisted self-correction processes characteristic of data-driven approaches to pronunciation learning (; ).

From a methodological perspective, the combination of music, gamification, and acoustic visualization makes it possible to construct a multimodal learning environment that integrates perception, production, and reflection.

The intervention follows an integrative pedagogical logic in which each component fulfills a complementary function. Music provides structured rhythmic and intonational input that models stress-timed organization; gamification, implemented through LyricsTraining, promotes repeated exposure and task persistence while reducing affective resistance to oral production; and acoustic visualization, introduced through illustrative use of Praat, supports metacognitive awareness by making otherwise abstract prosodic contours visually accessible. Together, these components create a perception–production–reflection cycle designed to facilitate the internalization of functional prosodic patterns.

From a causal perspective, music primarily strengthens auditory sensitivity to prominence through repeated exposure to rhythmically constrained input, which supports the perception of stress and intonational peaks. Gamified practice increases task repetition and persistence while reducing evaluative pressure, thereby expanding opportunities for procedural consolidation of suprasegmental routines. Acoustic visualization supports conscious noticing by making pitch movement perceptually salient, which can guide targeted self-monitoring during subsequent spoken practice. Under this model, perceptual attunement, sustained practice, and metacognitive noticing are expected to operate jointly and to produce gains in listener-perceived ease of understanding.

Such a design is particularly well suited to secondary education, where the effectiveness of an intervention depends not only on its theoretical grounding but also on its compatibility with the regular timetable, the official curriculum, and its capacity to sustain learner engagement under authentic classroom conditions. The combined assessment of objective acoustic measures, perceptual ratings of comprehensibility, and affective variables further offers a comprehensive account of the intervention's impact on oral language use.

This multimodal design is consistent with pedagogical frameworks that emphasize the meaningful integration of technology within instructional practice (), where digital tools are not treated as peripheral add-ons but are aligned with clearly defined learning objectives. In the present study, technological resources were embedded to support specific prosodic and metacognitive goals, rather than to introduce innovation for its own sake.

This study addresses a persistent gap between research on pronunciation in foreign languages and its systematic implementation in real secondary education contexts. Although the literature has consistently established the importance of suprasegmental features for speech comprehensibility, most classroom-based studies are characterized by short-term interventions, treatments relying on a single pedagogical resource, or evaluations focused exclusively on either acoustic or perceptual measures, without the simultaneous integration of affective dimensions or the examination of the temporal stability of observed effects.

In response to this shortcoming, the present study examines whether a multimodal pedagogical sequence, feasible within the constraints of the regular school timetable and grounded in the combined use of music, gamification, and acoustic visualization support, can generate functional prosodic changes that are perceptually relevant to native listeners and sustained beyond the immediate instructional period. The study makes a threefold contribution. First, it provides convergent evidence derived from instrumental acoustic analyses and blind perceptual ratings to assess prosodic development under ecologically valid classroom conditions. Second, it incorporates longitudinal follow-up data from an age group that remains underrepresented in pronunciation research, namely first-year students of compulsory secondary education. Third, it links observed prosodic changes to the development of key affective variables, thereby offering an empirical framework for understanding how motivation, self-efficacy, and anxiety interact with the development of oral competence in formal school settings.

2 Methodology

2.1 Study design

The research adopted a mixed-methods quasi-experimental design, integrating quantitative and qualitative data, with the aim of examining the effects of a pedagogical intervention based on music, gamification, and acoustic visualization on prosody, comprehensibility, and affective variables in the learning of English as a foreign language. The study followed an applied and ecological approach, as it was conducted entirely within the regular school timetable and without any modification to the school's curricular organization.

The choice of a quasi-experimental design was determined by the practical constraints of the school context, in which individual random assignment of learners is not feasible. Nevertheless, a series of selection and allocation procedures was implemented to minimize the main threats to internal validity typically associated with this type of design.

The participating school comprised nine first-year groups of compulsory secondary education, established by the school management at the beginning of the academic year according to general organizational rather than academic criteria. All groups displayed a balanced gender distribution, and all students had previously completed primary education across all subjects, with no differentiation based on prior academic attainment. From this population, two intact classes were randomly selected to take part in the study. This random selection was conducted at the level of pre-existing classes rather than individual learners, and no intervention by the research team took place in the internal composition of the classes, which had been constituted prior to the onset of the research.

Once the two classes had been selected, assignment to either the experimental or the control condition was also carried out at random at the class level. This cluster-level random assignment ensured that group allocation was not influenced by pre-existing learner characteristics or by instructional decisions, while acknowledging that individual randomization was not feasible in this educational setting. Initial equivalence between conditions was subsequently verified through statistical analysis of the pretest scores.

The design comprised three data collection points. A pretest was administered prior to the intervention in order to establish baseline measures for the variables under analysis. Following the completion of the intervention, an immediate post-test was conducted to assess short-term effects. Finally, a delayed post-test was administered three months later with the aim of examining the stability of learning outcomes in the absence of any additional targeted instruction during that period.

2.2 Educational context and participants

The study was conducted in a public lower secondary school located in a medium-sized urban setting, serving students from diverse sociocultural backgrounds and following the official curriculum for the teaching of English as a foreign language at the compulsory secondary level. The research was carried out within the regular school timetable, without introducing any changes to the general organization of the school or to the distribution of subjects.

The participants were first-year students of compulsory secondary education (N = 40), aged between 11 and 12 years. Twenty students were assigned to the experimental group (EG) and twenty to the control group (CG). The school does not implement ability grouping at this stage, and classes are formed without reference to prior academic performance. None of the participants required significant curricular adaptations in the foreign language area, nor did any present hearing impairments that could interfere with tasks involving oral perception and production.

In both groups, English instruction was delivered with the same number of weekly hours, the same coursebook, and the same general departmental syllabus. The teaching staff held appropriate subject-specific qualifications and had experience in secondary education. Mandatory curricular content, general assessment criteria, and the number of sessions devoted to the subject were kept constant, such that the only difference between conditions lay in the methodology used to address pronunciation and prosody.

It is also worth noting that both groups were taught by the same teacher throughout the duration of the study, which allowed variables related to teaching style, classroom management, and pedagogical interaction to be held constant. Furthermore, the intervention did not involve an increase in the amount of time allocated to English, but rather a redistribution of the instructional focus within the regular timetable and the official syllabus. Consequently, any differences observed between groups cannot be attributed to increased exposure time, but rather to the nature of the pedagogical practices implemented. This consideration further strengthens the internal validity of the design under authentic classroom conditions.

Prior to the commencement of the study, authorization was obtained from the school management team, and informed consent was secured from the students' families. The study was conducted in accordance with institutional and national ethical standards for research involving human participants. Data confidentiality and participant anonymity were guaranteed at all times.

2.3 Pedagogical intervention

The pedagogical intervention aimed to promote the development of English prosody through the integration of music, gamified technology, and visual support, in alignment with the theoretical framework underpinning the study. The intervention was implemented exclusively with the experimental group, while the control group continued to receive regular English instruction in accordance with the school's standard teaching practices.

The intervention was implemented over a period of six consecutive weeks, with a frequency of four weekly sessions of 55 min, integrated into the regular school timetable. The activities were embedded within the standard course syllabus and addressed compulsory curricular content through an alternative methodological approach that foregrounded prosodic work through the use of songs. All planned sessions were delivered as scheduled.

The core of the intervention was the recurrent use of English-language songs selected on the basis of their linguistic suitability, cultural familiarity for the learners, and prosodic potential. The songs displayed clear patterns of stress-based rhythm and intonation, as well as lexicon accessible to first-year secondary students. These materials were used as stimuli to support the integrated development of auditory perception and oral production, with particular attention to stress, rhythm, and fluency.

For listening practice, the digital platform LyricsTraining was employed to work with authentic songs through gamified activities based on active listening and the identification of words synchronized with the audio input. The use of this tool promoted repetition, heightened attention to rhythmic patterns, and learner self-regulation, while also reducing negative perceptions of error.

The sessions followed a relatively stable instructional sequence, comprising an initial global listening phase, focused listening activities targeting rhythmic and intonational patterns, and guided oral production tasks conducted both collectively and individually. These activities were introduced progressively, with an emphasis on fluency and participation rather than constant explicit correction. The intervention was designed to be readily replicable by secondary school teachers without the need for specialized technical resources or structural modifications to the existing curriculum.

In addition, brief moments of guided reflection were incorporated to draw learners' attention to specific prosodic features present in the songs under study. In selected sessions, acoustic visualization was also introduced through illustrative examples of intonation contours generated using the Praat software. This component served a formative and illustrative purpose only, and students did not interact directly with the software. Praat was used exclusively by the research team for the subsequent analysis of oral productions, while its classroom use was limited to visual demonstration. In this sense, acoustic visualization functioned primarily as structured metacognitive support rather than as a tool for direct student manipulation, providing a visual anchor for abstract suprasegmental features.

The two groups were taught in separate classrooms and followed distinct instructional sequences within their regular timetable. Access to the digital tools employed in the intervention (LyricsTraining and illustrative acoustic visualization) was restricted to the experimental group during scheduled sessions. These organizational measures were implemented to minimize the risk of cross-group contamination.

The overall temporal organization of the intervention is summarized in Table 1, which outlines the predominant activities by week, the resources employed, and the primary prosodic focus.

Table 1

WeekPredominant activityResources usedProsodic focus
1Initial listening explorationSongsRhythm and stress
2Guided rhythmic practiceSongs and gestural supportPhrasal stress
3Gamified practiceLyricsTrainingRepetition and segmentation
4Occasional visual supportPraatIntonation contours
5Communicative productionReading and guided speechFluency
6Final integrationMusic, LyricsTraining and PraatGlobal prosody

Overall schedule of the intervention in the experimental group.

During the same period, the control group covered the curricular content of the subject using the prescribed textbook and routine classroom activities. Pronunciation practice in this condition was limited to occasional repetition exercises and reading aloud, without the systematic use of songs, gamified platforms, or acoustic visualization.

2.4 Oral production tasks and selection of stimuli for acoustic analysis

The oral production tasks were designed with the aim of eliciting comparable and ecologically valid speech samples that would allow for the analysis of prosodic changes associated with the intervention. Speech data were collected at three time points, prior to the intervention, immediately after its completion, and three months later, in order to examine both immediate effects and their stability over time.

The tasks consisted of the oral production of short sentences selected from the songs used during the intervention, together with controlled utterances that are widely employed in prosodic research. The sentences displayed clearly identifiable prosodic patterns appropriate to the learners' proficiency level, which enabled systematic analysis of features such as stress timing, temporal organization of the utterance, and tonal variation.

All productions were spoken rather than sung, with the explicit aim of analyzing the transfer of the prosodic patterns practiced during the intervention to controlled spoken discourse. Recordings were carried out individually in a quiet room within the school, using the same recording device and following a standardized instruction protocol. The sentences were presented in the same fixed order for all participants and at all data collection points in order to ensure consistency across recordings.

The audio samples were stored in digital format and coded to ensure participant anonymity. They were subsequently segmented and analyzed using the Praat software, focusing on temporal and melodic parameters relevant to English prosody in a foreign language context. See Table 2 for the list of sentences used in the acoustic analysis.

Table 2

Produced sentenceProsodic features analyzedData collection point
She sells seashells by the seashoreRhythm, pauses, and stressPretest, post-test, delayed post-test
How much wood would a woodchuck chuckFluency and rhythmic variation
It's a beautiful day in the neighborhoodTonal contour
I can't believe how fast this year is flying byMelodic control
The quick brown fox jumps over the lazy dogRhythmic regularity

Sentences selected for acoustic analysis and corresponding data collection points.

2.5 Data collection instruments

To assess the impact of the intervention, a methodological triangulation was adopted that integrated acoustic, perceptual, and affective data collection instruments.

2.5.1 Instrumental acoustic analysis of f₀ range (Praat)

The oral production recordings constituted the basis for speech signal analysis using Praat software (version 6.4.07). Both temporal parameters, including total duration and articulation rate, and melodic parameters were examined. Particular emphasis was placed on the range of the fundamental frequency (f₀), normalized in semitones relative to 100 Hz. In order to ensure signal clarity in adolescent voices, the pitch range was set between 100 and 600 Hz, and tonal variability was operationalized as the difference between the 90th and 10th percentiles (P90–P10), following established methodological recommendations ().

The f₀ range (expressed in semitones) was selected as the central melodic metric due to its robustness in capturing pitch variability and its direct relevance to intonational expressiveness, which constituted the primary prosodic focus of the music-based intervention. As a global indicator of melodic expansion, f₀ range provides a reliable representation of changes in tonal dynamics in adolescent voices under classroom recording conditions. Although additional rhythm-related indices may offer further insight into temporal restructuring, f₀ range was prioritized because of its interpretability and stability in ecologically valid educational settings.

Acoustic measurements were obtained using a semi-automatic procedure and were subsequently reviewed by the research team to ensure consistency across samples.

2.5.2 Perceptual assessment of comprehensibility

Prosody and overall comprehensibility were assessed following a scoring protocol based on the multidimensional scale developed by . This evaluation framework, commonly referred to in applied contexts as POTS (Pronunciation Obstacle Detection and Scoring), allows for the quantification of the impact of prosodic deviations on speech clarity. For the purposes of the present study, descriptors from this scale were used to enable three native speakers of English, with prior experience in language teaching and pronunciation assessment, to evaluate, in a blind and randomized manner, the adequacy of participants' rhythm, stress, and intonation. Inter-rater reliability was examined prior to statistical analysis.

In line with the conceptual distinction outlined above, the use of the POTS descriptors in this study was intended to capture listener-perceived ease of understanding as mediated by prosodic features, rather than degree of accent or phonetic nativeness. Thus, the operationalization adopted corresponds directly to the construct of comprehensibility as defined in contemporary L2 speech research.

2.5.3 Affective variables

Three self-report questionnaires with established psychometric validation for the Spanish educational context were administered in order to assess affective variables associated with the learning of English pronunciation. Intrinsic motivation was measured using the interest and enjoyment subscale of the Intrinsic Motivation Inventory (IMI; ). As no version validated specifically for English learning in Spanish secondary education was available, the original items were translated and adapted into Spanish, taking as a reference the dimensions of intrinsic motivation in second language learning described by .

To ensure content validity and linguistic appropriateness for the participants' age range (11–12 years), the translated version of the IMI subscale was reviewed by an expert panel comprising two doctoral-level applied linguists and three secondary school teachers, one of whom held a doctorate. The adapted version was subsequently subjected to a back-translation procedure. The reliability of the adapted scale in the present study was excellent, yielding a Cronbach's alpha coefficient of.88.

Communicative self-efficacy was measured using a self-report scale adapted to the learning of English in compulsory secondary education. The items were designed on the basis of self-efficacy construct and its application to foreign language learning in previous research, such as the study by .

Pronunciation-related anxiety was assessed using the Foreign Language Classroom Anxiety Scale (FLCAS; ), employing the Spanish version validated by .

All questionnaires were administered at pretest and immediate post-test using five-point Likert scales. Across the three instruments, internal consistency was high, with Cronbach's alpha values ranging from.82 to.89.

2.5.4 Qualitative data

As a complementary source of evidence, semi-structured group interviews were conducted with students from the experimental group after completion of the intervention. The interview guide was designed to explore students' subjective perceptions of the meaningfulness of music-based activities, acoustic visualization support, and gamified dynamics within their learning process, as well as their perceived impact on oral participation and confidence.

The interviews were audio-recorded, transcribed verbatim, and anonymized prior to analysis. The qualitative data were analyzed using thematic analysis, following an inductive and interpretive approach aimed at identifying recurrent patterns in participants' perceptions. The analysis focused on capturing common trends and salient themes that could help contextualize and enrich the interpretation of the quantitative findings, rather than on the development of a formal categorical system.

2.6 Procedure

The study procedure was organized into three phases. During the initial phase, a pretest was administered in both groups. This phase included the oral production tasks and the administration of the affective variable questionnaires, following a standardized protocol.

Subsequently, the pedagogical intervention was implemented over a six-week period in the experimental group, while the control group continued to receive regular instruction. Throughout this phase, organizational conditions, curricular content, and the number of instructional sessions were held constant across both groups.

Following the completion of the intervention, an immediate posttest was conducted in both groups, replicating the oral production tasks and questionnaire administration procedures. Finally, a delayed posttest was carried out three months later, during which students once again produced the same recorded sentences as in the previous phases, without having received any additional instruction specifically targeting prosody during the intervening period.

All recordings and questionnaire data were stored and coded anonymously, ensuring data traceability and consistency in subsequent analyses.

2.7 Data analysis

Data analysis followed a mixed-methods approach, integrating quantitative and qualitative procedures. Acoustic data derived from the oral production recordings were analyzed using Praat software through systematic measurement of temporal and melodic parameters relevant to English prosody in a foreign language context.

The resulting values were subjected to descriptive and inferential statistical analyses. Given the repeated-measures design, repeated-measures analyses of variance were conducted to examine the effects of time, group, and the interaction between these factors. Where appropriate, post hoc comparisons were performed. The level of statistical significance was set at p < .05.

Perceptual ratings obtained through the POTS scale were analyzed following a similar procedure. Inter-rater reliability was examined, and scores were compared across groups and time points in order to track the development of prosody and comprehensibility.

Data from the affective variable questionnaires were analyzed using quantitative statistical techniques, comparing pretest and posttest scores across both groups.

Finally, qualitative data derived from the group interviews were analyzed thematically in order to identify recurrent patterns in learners' perceptions and to provide complementary interpretive support for the quantitative results.

2.8 Ethics statement

The study was conducted in accordance with the ethical guidelines of the Valencian Regional Ministry of Education. The research protocol was officially approved by the Secretaría Autonómica de Educación (Valencian Regional Government), with the identifier CSV: AERNMM9S:C6SF2Y23:N5UQ62DG, in a formal resolution dated November 24, 2023. Additionally, the study was approved by the school's governing board. Written informed consent was obtained from the parents or legal guardians of all participants, and written assent was provided by the students themselves.

3 Results

3.1 Quantitative results: effect of the intervention on prosody and comprehensibility

3.1.1 Initial equivalence and descriptive statistics

The first step of the analysis consisted of verifying initial equivalence between the EG and the CG. Independent-samples t tests were conducted on pretest scores across the five assessed dimensions (S1–S5). The results revealed no statistically significant differences in any of the variables analyzed (p > .05 in all cases), indicating that both groups started from a comparable level of prosodic and communicative competence. This initial equivalence allows subsequent changes observed in later measurements to be interpreted as attributable to the differential effect of the pedagogical intervention.

3.1.2 Repeated-measures analysis of variance (ANOVA)

Repeated-measures ANOVAs revealed significant Group × Time interaction effects across the main prosodic variables, indicating distinct developmental trajectories between groups. Highly significant interactions were found for intonation [S1: F(1, 38) = 171.00, p < .001, ηp2 = .82], stress [S2: F(1, 38) = 6.78, p = .013, ηp2 = .15], and rhythm and fluency [S3: F(1, 38) = 24.51, p < .001, ηp2 = .39], reflecting substantial improvements in the experimental group compared to the relative stability of the control group (see Table 3).

Table 3

DimensionGroupPretest (M SD)Post-test (M SD)Delayed post-test (M SD)
Intonation (S1)EG5.20 (0.8)7.90 (0.6)7.40 (0.7)
CG5.15 (0.9)5.30 (0.8)5.20 (0.9)
Stress (S2)EG5.40 (0.7)7.85 (0.5)7.35 (0.6)
CG5.30 (0.8)5.45 (0.7)5.35 (0.8)
Rhythm (S3)EG4.80 (1.1)7.60 (0.7)7.25 (0.8)
CG4.90 (1.0)5.10 (0.9)4.95 (1.0)
Phonemes (S4)EG5.05 (0.9)5.20 (0.8)5.10 (0.9)
CG5.00 (1.1)5.10 (1.0)5.05 (1.1)
Comprehensibility (S5)EG5.30 (0.8)7.70 (0.6)7.20 (0.7)
CG5.10 (0.9)5.20 (0.8)5.15 (0.9)

Mean scores and standard deviations for prosodic dimensions and overall comprehensibility (N = 40).

Scores were obtained using a 1–9 scale. M, mean; SD, standard deviation; EG, experimental group; CG, control group.

A post hoc power analysis was conducted using G*Power 3.1 to assess the adequacy of the sample size. Based on the observed interaction effect sizes (partial η2 ranging from.15 to.82) and a total sample of 40 participants, the statistical power (1–β) for the principal interaction effects exceeded.85. These values meet conventional thresholds for acceptable statistical sensitivity and further support the robustness of the reported findings.

3.1.3 Instrumental acoustic analysis of f₀ range (Praat)

A 2 × 3 mixed ANOVA on f₀ range revealed a significant Group × Time interaction [F(2, 76) = 18.34, p < .001, ηp2 = .32]. Post hoc comparisons with Bonferroni correction confirmed that while groups were equivalent at pretest (p > .05), the experimental group exhibited a significant increase in f₀ range at posttest (p < .001, d = 1.02), which remained significantly above baseline at the three-month follow-up (p < .01, d = 0.78). No significant variations were observed in the control group (p > .05). These results provide objective acoustic evidence of a sustained expansion in melodic range in the experimental group (see Table 4).

Table 4

Group (n = 20)Pretest M (SD)Post-test M (SD)Delayed post-test M (SD)
Experimental6.00 (2.1)8.10 (2.0)7.60 (2.0)
Control5.90 (2.1)6.30 (2.2)6.00 (2.2)

Descriptive statistics for f₀ range in semitones by group and assessment point (N = 40), calculated as the difference between the P90 and P10 percentiles using praat.

M denotes the mean and SD the standard deviation. The f₀ range was obtained as the difference between the 90th and 10th percentiles.

3.1.4 Durability of effects

The durability of the intervention effects was examined through a delayed post-test administered three months after the completion of the intervention, with the aim of assessing the stability of the gains observed in the immediate post-test. Overall, the results reveal a slight attenuation of performance at the delayed follow-up, a pattern commonly reported in educational research and reflective of the natural consolidation curve of learning in the absence of targeted reinforcement.

Despite this moderate decline, the EG maintained performance levels that were clearly higher than those recorded at pretest across all evaluated dimensions, indicating substantial retention of the intervention effects. In contrast, the CG exhibited a stable pattern across the three assessment points, with no evidence of sustained progress.

As illustrated in Figure 1, the temporal trajectory of the EG is characterized by a pronounced improvement following the intervention and relative stability at the delayed measurement, whereas the CG remains largely unchanged. Taken together, these results reinforce the interpretation that the effects observed were not limited to an immediate impact, but demonstrated medium-term persistence, in line with the instrumental acoustic evidence presented in the preceding section.

Figure 1

3.1.5 Association between acoustic expansion and perceptual gains

To further examine the relationship between instrumental and perceptual evidence, Pearson correlation analyses were conducted between individual gains in f₀ range (semitones) and improvements in perceived comprehensibility within the experimental group (n = 20). As shown in Table 4, mean f₀ range increased from 6.00 semitones (SD = 2.10) at pretest to 8.10 semitones (SD = 2.00) at post-test, yielding an average gain of 2.10 semitones. In parallel, comprehensibility scores increased from 5.30 (SD = 0.80) to 7.70 (SD = 0.60), with a mean gain of 2.40 points on the 1–9 scale (Table 3).

Correlation analyses revealed a significant positive association between melodic expansion and perceptual improvement, r = .59, p = .006. Participants who exhibited greater increases in f₀ range also tended to show larger gains in comprehensibility. This result suggests that pitch-range expansion was systematically associated with enhanced listener perception, reinforcing the convergence between acoustic and perceptual measures.

3.2 Qualitative findings

Qualitative data from group interviews and classroom observations corroborated the quantitative findings, particularly with regard to affective engagement and prosodic awareness. A primary theme was the reduction of affective barriers during oral production, as participants reported that the playful nature of music-based and gamified activities diminished evaluative pressure and encouraged risk-taking. As one student explained, “Singing makes it feel like a game… so I dare to try sounds without being afraid of being corrected” (GE1B). This emotional release was frequently associated with greater willingness to experiment with English prosody.

Regarding linguistic awareness, students consistently identified songs as tools that facilitated the internalization of rhythmic flow and intonation. Participants reported improved control over pausing and speech continuity, noting that musical input helped them reduce monotony and avoid inappropriate pauses. While both groups acknowledged initial difficulties with complex articulatory patterns, only the experimental group explicitly linked their progress to the engagement generated by digital platforms such as LyricsTraining, which promoted voluntary practice beyond the classroom.

Finally, the intervention fostered a positive social environment characterized by increased peer support and reduced embarrassment during oral tasks. Although both groups valued collaboration, students in the experimental group reported feeling more comfortable participating orally. Despite differences in musical preferences, the pedagogical materials were generally perceived as meaningful and effective.

3.3 Results for motivation, self-efficacy and anxiety

Mixed ANOVAs on affective variables revealed significant Group × Time interaction effects across all dimensions. The experimental group showed significant increases in intrinsic motivation [F(1, 38) = 36.48, p < .001, ηp2 = .49, d = 1.20] and communicative self-efficacy [F(1, 38) = 31.60, p < .001, ηp2 = .45, d = 1.10], whereas no statistically significant changes were observed in the control group (p > .05). In addition, a highly significant reduction in pronunciation-related anxiety was recorded exclusively in the experimental group [F(1, 38) = 27.92, p < .001, ηp2 = .42, d = 1.05], while the control group remained stable over time (p > .05).

These results indicate that the intervention was associated with systematic changes in affective variables closely related to oral production during adolescence (see Table 5).

Table 5

VariableGroup (N = 20)Pretest M (SD)Posttest M (SD)
Intrinsic motivationExperimental3.50 (0.62)4.40 (0.55)
Control3.48 (0.60)3.55 (0.58)
Communicative self-efficacyExperimental3.40 (0.65)4.30 (0.57)
Control3.42 (0.63)3.50 (0.61)
Pronunciation anxietyExperimental3.70 (0.68)2.80 (0.60)
Control3.65 (0.66)3.55 (0.64)

Descriptive statistics for affective variables by group and assessment point (N = 40).

Scores were obtained using five-point Likert scales. For the anxiety variable, lower scores indicate lower perceived anxiety.

A Pearson correlation analysis was conducted to explore the relationship between affective development and linguistic outcomes. Gains in communicative self-efficacy were positively correlated with improvements in perceived comprehensibility (r = .48, p < .01). This finding supports the theoretical proposition that increased confidence in oral production is associated with enhanced prosodic performance, highlighting the interplay between affective and linguistic development.

3.3.1 Intrinsic motivation

The analysis revealed a significant main effect of Time, F(1, 38) = 42.15, p < .001, ηp2 = .53, as well as a significant Group × Time interaction, F(1, 38) = 36.48, p < .001, ηp2 = .49. Post hoc comparisons indicated that the experimental group showed a significant increase in intrinsic motivation from pretest to posttest (p < .001, d = 1.20), whereas no statistically significant changes were observed in the control group (p > .05).

3.3.2 Communicative self-efficacy

A significant main effect of Time was observed, F(1, 38) = 38.72, p < .001, ηp2 = .50, together with a significant Group × Time interaction, F(1, 38) = 31.60, p < .001, ηp2 = .45. The experimental group exhibited a significant increase in perceived communicative self-efficacy between pretest and posttest (p < .001, d = 1.10), while no significant changes emerged in the control group (p > .05).

3.3.3 Pronunciation anxiety

The analysis revealed a significant main effect of Time, F(1, 38) = 29.84, p < .001, ηp2 = .44, as well as a significant Group × Time interaction, F(1, 38) = 27.92, p < .001, ηp2 = .42. Post hoc analyses confirmed that the experimental group experienced a significant reduction in pronunciation-related anxiety from pretest to posttest (p < .001, d = 1.05), whereas the control group remained stable over time (p > .05).

4 Discussion

Before interpreting the observed effects, it is necessary to consider issues related to causal attribution in quasi-experimental studies conducted in authentic educational contexts. The intervention incorporated elements perceived by students as more motivating and dynamic, which may have generated a novelty effect and a higher level of initial engagement in the experimental group. However, rather than constituting a confounding variable, this increase in engagement forms an integral part of the proposed pedagogical design, whose explicit aim is to create emotionally and attentively supportive conditions for sustained prosodic practice during adolescence.

From this perspective, the results should not be interpreted as the isolated effect of a single pedagogical resource, but rather as the outcome of an intentional reorganization of classroom practices within the available instructional time, oriented towards prioritizing functional prosody and active learner participation. Although it is not possible to fully isolate the specific contribution of each component or to entirely rule out the influence of a novelty effect, the partial stability of the results observed in the delayed post-test, together with their convergence across acoustic and perceptual measures, suggests that the observed improvements cannot be reduced to a transient surge of enthusiasm, but instead reflect functional changes in oral production.

The findings of the present study suggest that a multimodal pedagogical intervention grounded in music, gamified technology and acoustic visualization support is associated with meaningful improvements in prosody and comprehensibility in English as a foreign language among lower secondary students. In line with previous research highlighting the central role of prosody as a predictor of comprehensibility over and above segmental accuracy, the quantitative analyses revealed robust intervention effects on key variables such as intonation, stress, rhythm and overall comprehensibility, while no differential gains were observed in the articulation of isolated phonemes (; ; ). This dissociation reinforces the pedagogical relevance of approaches that prioritize the prosodic organization of discourse over traditional models centered on segmental correction, particularly in educational contexts where exposure time to the target language is limited.

The convergence of statistical results, native listeners' perceptual ratings and illustrative acoustic analysis using Praat provides coherent evidence that the observed gains are not merely superficial, but rather reflect a functional reorganization of the rhythmic and melodic patterning of speech. In particular, the shift from a pattern approximating syllable isochrony towards a clearer hierarchical alternation between stressed and unstressed syllables, together with greater tonal continuity, aligns with established accounts of prosodic development in learners of English as a foreign language (; ). The fact that these changes emerged within an ecologically valid classroom context and were partially maintained at the delayed post-test further strengthens the applied value of the study and its potential relevance for routine teaching practice.

4.1 Effects of the intervention on learners' prosodic development

The results confirm that deliberate prosodic instruction, implemented through multimodal resources, can generate structural changes in oral production that extend beyond the scope of conventional instruction. The most salient outcome is the significant expansion of the f₀ range in the experimental group, which reached a mean increase of 2.1 semitones. This increase in melodic range does not represent a superficial modification, but rather reflects a transition away from the tonal monotony associated with the transfer of Spanish syllable-timed rhythm towards a more dynamic and functionally appropriate intonational architecture in English (; ).

This melodic reorganization is consistent with Patel's OPERA hypothesis (2011), according to which the heightened demands for auditory precision imposed by musical training refine neural networks shared with language and promote more stable encoding of tonal peaks. While the control group maintained a restricted melodic range (M = 6.0 ST), the experimental group displayed greater tonal expressiveness which, according to native listeners' perceptual ratings using the POTS scale, facilitated discourse segmentation and the identification of meaning units.

With regard to rhythm, the improvement observed in variable S3, together with its reflection in reduced utterance duration, points to increased efficiency in speech motor planning. The reduction of compensatory lengthening in unstressed positions suggests that learners had begun to automatize vowel reduction processes, one of the principal challenges faced by Spanish-speaking learners of English (). This shift towards a rhythmic pattern more closely aligned with stress timing is consistent with account, which links oral fluency to the release of attentional resources resulting from the automatization of suprasegmental features.

In contrast to these suprasegmental gains, no differential improvement was recorded in the articulation of specific phonemes (variable S4). This dissociation between segmental and suprasegmental development may reflect the top-down orientation of the intervention, which prioritized global rhythmic flow, intonational contour, and fluency over isolated phoneme-level correction. Learners appear to have reorganized the suprasegmental gestalt of speech before refining individual articulatory targets, a trajectory consistent with communicative approaches that foreground prosodic coherence as a foundation for later segmental refinement (; ). This pattern further supports the view that global comprehensibility is more strongly determined by coherent prosodic organization than by isolated phonetic accuracy, underscoring the pedagogical value of prioritizing functional prosody over approximation to a nativist ideal.

4.2 Effects of the intervention on overall speech comprehensibility

The second objective of the study was to determine whether the acoustic improvements observed in prosody translated into perceptible communicative benefits. The results derived from native listener evaluations indicate an increase of 2.4 points on the comprehensibility scale for the experimental group. This gain is not only statistically significant but also pedagogically and clinically meaningful, as reflected in its large effect size (d = 1.22).

This finding is particularly salient given that it emerged in the absence of improvements in segmental accuracy, as captured by variable S4. The convergence of these results supports the view that comprehensibility does not depend directly on isolated phonetic correctness, but rather on the clarity with which speakers organize units of meaning through suprasegmental features (). From a neurocognitive perspective, the shift towards a pattern more closely resembling stress timing, together with enhanced melodic continuity, contributed to a reduction in processing load for native listeners. By providing a more predictable rhythmic and tonal structure, the speech produced by the experimental group facilitated discourse segmentation and the identification of salient lexical items with reduced cognitive effort ().

In contrast, the control group maintained virtually stable comprehensibility scores, with values ranging between 5.1 and 5.2. This pattern suggests that traditional instruction, often centered on reactive correction of phonetic errors, fails to reach the threshold required to enhance the ease with which continuous speech is understood. While the control group continued to produce speech characterized by tonal fragmentation and syllable-timed rhythm, features that compel native listeners to engage in word-by-word processing, the experimental group developed a prosodic organization that directed listeners' attention towards the informational hierarchy of the message.

These findings carry direct implications for the teaching of English in compulsory secondary education, where instructional time is limited. The intervention demonstrates that prioritizing functional prosody over segmental precision can yield higher levels of communicative effectiveness within a relatively short timeframe, in this case six weeks. This supports the integration of multimodal pedagogical approaches within the regular secondary curriculum.

4.3 Development of motivation, communicative self-efficacy and pronunciation anxiety

The third objective of the study examined the trajectory of affective variables, based on the premise that prosodic competence is closely mediated by learners' emotional states. The results indicate that the intervention brought about significant changes in the affective profile of the experimental group, with increases of more than one point on the intrinsic motivation and self-efficacy scales, alongside a marked reduction in anxiety (d = 1.05). By contrast, the control group maintained largely stable levels of anxiety and motivation, suggesting that traditional instruction does not effectively mitigate the psychological barriers associated with oral production during adolescence.

The reduction in communicative anxiety observed in the experimental group is central to interpreting the prosodic gains reported in this study. In line with affective filter hypothesis, a learning environment perceived as safe and playful facilitates the lowering of defenses associated with social evaluation. In this regard, the use of singing and gamified dynamics transformed phonetic practice, often experienced by adolescents as a high-risk activity for public self-image, into a collective experience with a reduced emotional load. The normalization of error through immediate feedback provided by platforms such as LyricsTraining fostered a more open attitude towards prosodic exploration, an outcome that is difficult to achieve through approaches based on mechanical repetition ().

Similarly, the increase in communicative self-efficacy (d = 1.10) suggests that learners in the experimental group developed greater confidence in their ability to manage suprasegmental aspects of speech. From the perspective of social cognitive theory, such perceptions of competence exert a direct influence on sustained effort and persistence when facing demanding tasks. In the present study, illustrative acoustic visualization through Praat functioned as structured metacognitive support, offering learners a visual representation of prosodic development, while successful engagement with gamified challenges reinforced sustained practice and progress. Together, these elements contributed to strengthening learners' perceptions of communicative competence. This enhanced confidence is likely to promote more frequent and attentive practice of intonation and rhythm, generating a reciprocal relationship between perceived capability and linguistic performance ().

Although the positive correlation observed between gains in communicative self-efficacy and improvements in perceived comprehensibility (r = .48, p < .01) suggests a functional association between affective development and prosodic performance, the present sample size does not permit stable estimation of a formal mediation model. With only twenty participants in the experimental group, indirect effect estimates would be statistically underpowered and highly sensitive to sampling variability. For this reason, the current study interprets the affective–prosodic link cautiously, relying on convergent interaction effects and correlational evidence rather than on causal mediation claims. Future research with larger samples and temporally staggered measurements would be required to determine whether affective gains statistically mediate improvements in prosodic organization and listener-perceived ease of understanding.

Taken together, the coherence between affective outcomes and acoustic data supports the effectiveness of the multimodal pedagogical model adopted in this study. The intervention not only offered technical resources for the development of prosody, but also created emotionally supportive conditions conducive to its effective application. These findings highlight the importance of systematically integrating the socio-affective dimension into the design of pronunciation teaching approaches, particularly at educational stages characterized by heightened sensitivity to external evaluation, such as compulsory secondary education.

4.4 Durability of effects and stability of prosodic learning

Interpreting the durability of effects in educational intervention studies requires caution, especially when working with moderate sample sizes and without strict experimental control. In the present study, the delayed post-test does not allow for claims of full stabilization of prosodic change. Nevertheless, it provides evidence of partial retention in the absence of additional targeted instruction, a finding that is informative from an applied pedagogical perspective.

The fourth objective of the study examined the temporal stability of the effects achieved following the completion of targeted instruction. Results from the delayed post-test indicate that a substantial proportion of the improvements in the experimental group's tonal range were maintained three months after the intervention. Although a slight regression was observed, with mean values decreasing from 8.1 ST in the immediate post-test to 7.6 ST at the delayed measurement, performance remained significantly above the baseline level of 6.0 ST (p < .01, d = 0.78). This pattern of retention suggests that the prosodic changes observed cannot be explained solely as a transient novelty effect, but rather point towards signs of functional reorganization of melodic patterns.

The maintenance of these gains is particularly noteworthy given that prosodic acquisition is often characterized by strong resistance to stable change in the absence of continued exposure (; ). The observed stability may be interpreted in light of the multimodal and multisensory nature of the intervention. The combination of musical input, embodied activities and acoustic visualization appears to have fostered deeper and more redundant encoding of rhythmic and melodic patterns. This form of distributed learning, supported by multiple perceptual channels, facilitates the consolidation of more robust procedural memory traces, allowing prosodic patterns to remain accessible even in the absence of subsequent intensive practice (; ).

The modest decline observed at the delayed post-test is consistent with previous research indicating that prosodic development in a foreign language does not follow a linear trajectory. Partial regression in certain parameters suggests that fine-grained control of intonation tends to weaken over time if not reinforced through frequent use or sustained motivational conditions (). This pattern underscores the importance of moving away from isolated instructional interventions and of integrating prosodic practice recurrently within the regular lower secondary curriculum.

Taken together, the evidence of partial retention observed supports the applied potential of the multimodal pedagogical model. The intervention yielded immediate effects on comprehensibility and tonal variability while simultaneously fostering the development of prosodic competence with the capacity for medium-term maintenance. From a pedagogical perspective, these findings support the systematic integration of gamification strategies and acoustic support in the classroom in order to facilitate progressive, functional and sustainable consolidation of oral skills.

Beyond corroborating previously documented findings on the role of prosody in comprehensibility, the primary contribution of this study lies in the articulation of an integrated pedagogical model that connects prosodic development, affective engagement and curricular feasibility in secondary education. Unlike studies focusing on isolated tools or techniques, the present results suggest that the observed gains emerge from the joint reorganization of classroom practices, in which attention to prosodic form, reduction of communicative anxiety and sustained learner participation operate in an interdependent manner.

From this perspective, the intervention should not be understood as a set of resources to be transferred mechanically, but rather as an instructional design logic that prioritizes functional prosody within the temporal and organizational constraints of the regular curriculum. This approach advances understanding of how pronunciation development can be promoted effectively and sustainably in formal school contexts, providing an empirical framework that extends beyond the simple replication of previously established effects.

4.5 Limitations and future directions

Despite the robustness of the findings, the present study is subject to a number of limitations that should be taken into account when interpreting and generalizing the results. First, the quasi-experimental design was implemented with intact groups within a single educational institution. Although initial equivalence was established through pretest measures and appropriate statistical analyses were conducted, the absence of individual random assignment and the moderate sample size (N = 40) call for caution when extrapolating the findings to socioculturally diverse contexts.

External validity is also constrained by the linguistic background of the sample, as all participants shared Spanish as their first language. Spanish–English rhythmic distance is often described in terms of syllable timing vs. stress timing, and this specific contrast may shape both the baseline difficulty of prosodic targets and the extent to which music-based rhythmic input facilitates change. Learners with first languages that are stress-timed or mora-timed may show different developmental trajectories, and the same intervention may yield distinct effects in intonation, rhythm, and perceived ease of understanding. Replication with typologically diverse L1 groups would therefore be necessary to more precisely determine the cross-linguistic scope of the proposed model.

Furthermore, although both groups were taught by the same teacher in order to control for instructional style and classroom management variables, it is not possible to entirely rule out the influence of teacher expectancy effects. The instructor was aware of the intervention condition, which may have unintentionally shaped subtle aspects of feedback, encouragement, or interactional dynamics. Nevertheless, several design features reduce this risk, including identical curricular content, equivalent instructional time, blind perceptual ratings by external evaluators, and the use of objective acoustic measures. Future research could incorporate multiple instructors or partially blinded instructional procedures to further minimize expectancy-related influences.

Moreover, affective variables were measured only at pretest and immediate post-test. The absence of delayed affective data limits the extent to which conclusions can be drawn regarding the long-term stability of motivational and anxiety-related changes. Future research should incorporate longitudinal affective measures to determine whether psychological gains parallel the durability observed in prosodic outcomes.

The internal consistency of the adapted IMI was high (α = .88). Although a CFA was considered to verify the factor structure, the sample size (N = 40) was not adequate for robust model estimation and stable fit indices according to established psychometric criteria (e.g., ; ). Consequently, psychometric support for the instrument in the present study derives from its strong internal reliability and its established use within research grounded in Self-Determination Theory in educational and second language learning contexts.

Second, the acoustic analysis focused on controlled utterances in order to ensure the comparability of data obtained through Praat. While this approach is necessary to guarantee instrumental rigor, it does not allow for precise determination of the extent to which prosodic gains transfer to spontaneous or unplanned speech, where cognitive load is higher. Future research should therefore incorporate free production tasks that enable assessment of transfer to more naturalistic communicative contexts.

Third, the multimodal nature of the intervention, which simultaneously integrated music, gamification and acoustic visualization, represents a pedagogical strength but also an analytical limitation. The adopted design does not permit precise discrimination of the specific contribution of each component to the observed improvements. It remains unclear whether the effects are primarily attributable to musical auditory training, to the metacognitive support provided by acoustic visualization, or to the motivational component associated with gamification. In this regard, future studies could employ factorial designs to examine the relative weight of each tool in the reorganization of prosodic patterns.

Finally, although the three-month follow-up provides valuable information regarding the temporal stability of the effects, extending the time frame to six or twelve months would be advisable in order to assess longer-term maintenance.

It should also be acknowledged that extended longitudinal tracking beyond this interval was not feasible within the structural organization of compulsory secondary education. In many educational systems, including the context in which this study was conducted, student groups are routinely reorganized at the end of each academic year. This widespread institutional practice prevents the maintenance of intact cohorts over successive years and limits the possibility of implementing longer-term follow-up assessments under comparable instructional conditions.

Further lines of research could also explore whether the development of prosodic competence is associated with gains in listening comprehension or with changes in learners' construction of linguistic identity as competent users of English.

5 Conclusion

This study demonstrates that a multimodal intervention integrating music, gamification, and acoustic visualization significantly improves prosody and speech comprehensibility in secondary EFL learners. Convergent evidence from acoustic and perceptual measures confirms that these gains reflect a functional reorganization of rhythm and intonation—key factors in L2 comprehensibility—rather than superficial changes.

Notably, the results show that significant improvements in comprehensibility can be achieved without parallel gains in segmental accuracy. This reinforces the pedagogical value of prioritizing functional prosody over traditional models focused exclusively on isolated phonetic correction. Specifically, the intervention successfully shifted learners away from syllable-timed patterns, promoted vowel reduction, and enhanced melodic control, with gains sustained three months post-instruction.

Furthermore, the findings underscore the importance of the affective dimension in pronunciation instruction. Increased motivation and self-efficacy, alongside reduced anxiety, created a safe learning environment conducive to sustained oral practice and long-term retention.

From an applied perspective, this research supports the systematic integration of music- and technology-based tools into secondary curricula to address oral skills, an area traditionally marginalized in formal settings. Ultimately, this study provides an ecologically valid model for promoting clearer, more functional, and sustainable oral competence in the EFL classroom.

Statements

Data availability statement

The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.

Ethics statement

The studies involving humans were approved by Secretaría Autonómica de Educación (Valencian Regional Government), Generalitat Valenciana (Reference: AERNMM9S:C6SF2Y23:N5UQ62DG). The studies were conducted in accordance with the local legislation and institutional requirements. Written informed consent for participation in this study was provided by the participants' legal guardians/next of kin. Written informed consent was obtained from the individual(s), and minor(s)' legal guardian/next of kin, for the publication of any potentially identifiable images or data included in this article.

Author contributions

VC-P: Project administration, Writing – original draft, Data curation, Methodology, Investigation, Conceptualization, Writing – review & editing. RE-F: Methodology, Writing – review & editing, Investigation, Data curation, Writing – original draft, Resources. MF-M: Writing – review & editing, Validation, Formal analysis, Methodology, Writing – original draft, Supervision. JE-F: Writing – review & editing, Conceptualization, Supervision, Writing – original draft, Software, Methodology, Formal analysis.

Funding

The author(s) declared that financial support was not received for this work and/or its publication.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

References

Summary

Keywords

acoustic visualization, educational psychology, EFL prosody, gamification, music-based learning, pronunciation anxiety, secondary education, speech comprehensibility

Citation

Chust-Pérez V, Esteve-Faubel RP, Fernández-Morante MC and Esteve-Faubel JM (2026) Integrating music, gamification, and acoustic visualization enhances prosody and comprehensibility in secondary EFL classrooms. Front. Educ. 11:1788875. doi: 10.3389/feduc.2026.1788875

Received

15 January 2026

Revised

01 March 2026

Accepted

06 March 2026

Published

25 March 2026

Volume

11 - 2026

Edited by

Ana Alexandra Silva, University of Evora, Portugal

Reviewed by

Serafeim A. Triantafyllou, Aristotle University of Thessaloniki, Greece

Lidia Cardoso, Federal University of Ceara, Brazil

Updates

Copyright

*Correspondence: José María Esteve-Faubel

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics