Abstract
One of the primary questions of second language (L2) acquisition research is how a new sound category is formed to allow for an L2 contrast that does not exist in the learner's first language (L1). Most models rely crucially on perceived (dis)similarities between L1 and L2 sounds, but a precise definition of what constitutes “similarity” has long proven elusive. The current study proposes that perceived cross-linguistic similarities are based on feature-level representations, not segmental categories. We investigate how L1 Japanese listeners learn to establish a new category for L2 American English /æ/ through a perception experiment and computational, phonological modeling. Our experimental results reveal that intermediate-level Japanese learners of English perceive /æ/ as an unusually fronted deviant of Japanese /a/. We implemented two versions of the Second Language Linguistic Perception (L2LP) model with Stochastic Optimality Theory—one mapping acoustic cues to segmental categories and another to features—and compared their simulated learning results to the experimental results. The segmental model was theoretically inadequate as it was unable explain how L1 Japanese listeners notice the deviance of /æ/ from /a/ in the first place, and was also practically implausible because the predicted overall perception patterns were too native English-like compared to real learners' perception. The featural model, however, showed that the deviance of /æ/ could be perceived due to an ill-formed combination of height and backness features, namely */low, front/. The featural model, therefore, reflected the experimental results more closely, where a new category was formed for /æ/ but not for other L2 vowels /ɛ/, /ʌ/, and /ɑ/, which although acoustically deviate from L1 /e/, /a/, and /o/, are nonetheless featurally well-formed in L1 Japanese, namely /mid, front/, /low, central/, and /mid, back/. The benefits of a feature-based approach for L2LP and other L2 models, as well as future directions for extending the approach, are discussed.
1 Introduction
Second language (L2) learners often encounter a “new” sound that does not exist in their first language (L1). Establishing a phonological representation for such new sounds is essential to L2 learning, because otherwise the lexical distinctions denoted by the phonological contrast cannot be made for successful communication. Various models have been proposed to explain how a new sound category may develop in the learner's mind, with most models focusing on the cross-linguistic perceptual relationships between L1 and L2 sounds, although the exact underlying mechanism remains to be elucidated. In this study, we propose that the process of L2 category formation can be better explained by assuming feature-level representations as the fundamental unit of perception, rather than segmental categories.1 To this end, we compare two versions of formal modeling, i.e., segment- and feature-based, of how L1 Japanese listeners form a new category for L2 American English (AmE) /æ/2 by implementing the theoretical predictions of the Second Language Linguistic Perception (L2LP) model (Escudero, ; van Leussen and Escudero, 2015; Escudero and Yazawa, ) with a computational-phonological approach of Stochastic Optimality Theory (StOT; Boersma, ).
In the field of L2 speech perception research, two models have been particularly dominant over the last few decades (Chen and Chang, ): the Speech Learning Model (SLM; Flege, ; Flege and Bohn, ) and the Perceptual Assimilation Model (PAM; Best, ; Best and Tyler, ). According to SLM, learners can form a new category for an L2 sound if they discern its phonetic difference(s) from the closest L1 category and, if not, a single composite category will be used to process both L1 and L2 sounds. The likelihood of category formation therefore depends primarily on the perceived cross-linguistic phonetic dissimilarity, but other factors such as the quantity and quality of L2 input obtained in meaningful conversations are said to be also relevant. PAM agrees with SLM in that perceived cross-linguistic dissimilarity guides category formation, but with the caveat that assimilation occurs not only at the phonetic level but also at the phonological or lexical-functional level. For example, if there are many minimal pairs involving an L2 contrast that assimilates phonetically to a single L1 category, the increased communicative pressure can lead to the formation of a new phonological category to allow for distinct phonological representations of these lexical items (e.g., AmE [ʌ] and [ɑ] both being assimilated to Japanese [a], but AmE nut [nʌt] and not [nɑt] leading to Japanese /na1t/ and /na2t/, where /a1/ and /a2/ are distinct phonological categories that occupy phonetically overlapping but distinct parts of a single L1 category). While the predictions of both models have been supported by numerous studies, there is one fundamental issue that remains to be resolved: It is unclear on what basis categorical similarity should be defined. In the words of Best and Tyler (, p.26), “one issue [...] has not yet received adequate treatment in any model of nonnative or L2 speech perception: How listeners identify nonnative phones as equivalent to L1 phones, and the level(s) at which this occurs.” Over a decade later, Flege and Bohn (, p.31) restated the unresolved problem: “It remains to be determined how best to measure cross-language phonetic dissimilarity. The importance of doing so is widely accepted but a standard measurement procedure has not yet emerged.”
To illustrate this elusive goal with a concrete example, consider our case of L1 Japanese listeners learning /æ/ and adjacent vowels in L2 AmE. In cross-linguistic categorization experiments, Strange et al. (1998) found that the AmE vowel was perceived as a very poor exemplar of Japanese /a(a)/,3 receiving the lowest mean goodness-of-fit rating (two out of seven) among all AmE vowel categories, while spectrally adjacent /ɛ/, /ʌ/, and /ɑ/ received higher ratings as Japanese /e/ (four out of seven), /a/ (four out of seven), and /a(a)/4 (six out of seven), respectively. Duration-based categorization of AmE vowels as Japanese long and short vowels was observed when the stimuli were embedded in a carrier sentence, but not when they were presented in isolation. Shinohara et al. (2019) further found that the category goodness of synthetic vowel stimuli as Japanese /a/ deteriorated as the second formant (F2) frequency was increased. These studies suggest that AmE /æ/ is perceptually dissimilar from L1 Japanese /a(a)/, presumably in terms of F2 but possibly in conjunction with other cues such as the first formant (F1) frequency and duration, and is thus subject to new category formation according to SLM and PAM. L2 perception studies on Japanese listeners also showed that the AmE /æ/-/ʌ/ contrast was more discriminable than the /ɑ/-/ʌ/ contrast (Hisagi et al., ; Shafer et al., 2021; Shinohara et al., 2022) and that AmE /æ/ was identified with higher accuracy than /ʌ/ or /ɑ/ (Lambacher et al., ). These results imply that Japanese listeners perceptually distinguish AmE /æ/ from AmE /ʌ/ and /ɑ/, which themselves are assimilated to Japanese /a(a)/. However, it remains unclear why only AmE /æ/ would be perceptually distinct in the first place. Spectral distance between the L1 and L2 categories in Figure 1, which shows the production of Japanese and AmE vowels by four native speakers of each language (Nishi et al., ), does not seem to predict perceived category goodness very well. For example, the figure shows that AmE /ʌ/ is spectrally closer than AmE /ɑ/ to Japanese /a(a)/, but it was the latter AmE vowel that was judged to be a better fit in Strange et al. (1998). AmE /ɛ/ is also quite far from Japanese /e/ in spectral distance, but its perceived category goodness as Japanese /e/ was nonetheless as high as that of AmE /ʌ/ as Japanese /a/. One may attribute this pattern to duration given the duration-based categorization in Strange et al. (1998), but the actual duration of the target vowels (Figure 2) does not seem to provide useful clues, either. For example, since AmE /ɛ/ and /ʌ/ have almost identical duration values, the former is acoustically more distant from the L1 categories after all, despite both AmE vowels receiving equal goodness ratings. This brings us back to the question: How is L1-L2 perceptual dissimilarity determined?
Figure 1
Figure 2

Average duration of relevant AmE and Japanese vowels. Adapted with permission from Nishi et al. (
The L2LP model approaches L2 category formation from a different perspective. While the model shares with SLM and PAM the view that perceptual learning is both auditory- and meaning-driven, it is unique in assuming the interplay of multiple levels of linguistic representations. Although many L2LP studies have focused on the perceptual mapping of acoustic cues onto segmental representations, some have also incorporated feature-level representations, which may be useful for modeling the perception of new L2 sounds. For example, Escudero and Boersma (
The incentive for feature-based modeling is not only theoretically grounded but also empirically motivated, as emerging evidence suggests the involvement of features in L1 and L2 perception. With respect to native perception, Scharinger et al. (2011) used magnetoencephalography (MEG) to map the entire Turkish vowel space onto cortical locations and found that dipole locations could be structured in terms of features (height, backness, and roundedness) rather than raw acoustic cues (F1, F2, and F3). Mesgarani et al. (
The remainder of this paper is organized as follows. First, in Section 2, we present a forced-choice perception experiment that investigates the use of spectral and temporal cues in the perception of the L1 Japanese and L2 AmE vowels of interest. This is intended to complement the previous studies, which did not investigate potential effects of F1 and duration cues. Section 3 then presents a formal computational modeling of new L2 category formation within the L2LP framework. Two versions of StOT-based simulations are compared, namely segmental and featural, to evaluate which better explains and replicates the experimental results. We then discuss the experimental and computational results together in Section 4, addressing the implications of feature-based modeling for L2 speech perception models (i.e., L2LP, SLM, and PAM) as well as the directions for future research. Finally, Section 5 draws the conclusion.
2 Experiment
The perception experiment reported in this section was designed to investigate how L2 AmE /æ/ and adjacent vowels are perceived in relation to L1 Japanese vowels based on three acoustic cues (F1, F2, and duration), to help model the category formation and cross-linguistic assimilation processes. Following our previous study (Yazawa et al., 2020), the experiment manipulates the ambient language context to elicit L1- and L2-specific perception modes without changing the relevant acoustic properties of the stimuli, as detailed below.
2.1 Participants
Thirty-six native Japanese listeners (22 male, 14 female) participated in the experiment. They were undergraduate or graduate students at Waseda University, Tokyo, Japan, between the ages of 18 and 35 (mean = 21.25, standard deviation = 2.97). All participants had received six years of compulsory English language education in Japanese secondary schools (from ages 13 to 18), which focused primarily on reading and grammar. They had also received some additional English instruction during college, the quality and quantity of which varied according to the courses they were enrolled in. None of the participants had spent more than a total of three months outside of Japan. TOEIC was the most common standardized test of English proficiency taken by the participants (n = 18), with a mean score of 688 (i.e., intermediate level). All participants reported normal hearing.
2.2 Stimuli
Two sets of stimuli—“Japanese” (JP) and “English” (EN)—were prepared. Both had the same phonetic form [bVs], with the spectral and temporal properties of the vowel varying in an identical manner. The JP stimuli were created from a natural token of the Japanese loanword baasu /baasu/ “birth,” as produced by a male native Japanese speaker from Tokyo, Japan. The token was phonetically realized as [ba:s] because Japanese /u/ can devoice or delete word-finally (Shaw and Kawahara, 2017; Whang and Yazawa, 2023). The EN stimuli, on the other hand, were created from a natural token of the English word bus /bʌs/, as produced by a male native AmE speaker from Minnesota, United States. For both tokens, the F1, F2, and duration of the vowel were manipulated with STRAIGHT (Kawahara,
2.3 Procedure
The experiment included two sessions—again “Japanese” (JP) and “English” (EN)—using the JP and EN stimuli, respectively. In order to elicit language-specific perception modes across sessions, all instructions, both oral and written, were given only in the language of the session. The two sessions were consecutive, and the session order was counterbalanced across participants to control for order effects; 18 participants (11 male, seven female) attended the EN session first, while the other 18 (11 male, seven female) attended the JP session first. In the JP session, participants were first presented with each of the 64 JP stimuli in random order and then chose one of the following four words that best matched what they had heard: beesu /beesu/ “base,” besu /besu/ “Bess,” baasu /baasu/ “birth,” and basu /basu/ “bus.” The choices are all existing loanwords in Japanese and were written in katakana orthography. Participants were instructed that they were not required to use all of the four choices. The block of 64 trials was repeated four times, with a short break in between, giving a total of 256 (64 × 4) trials for the session. The EN session followed a similar a procedure, where participants categorized the randomized 64 EN stimuli as the following four real English words (though there was no requirement to use all choices): Bess (/bɛs/), bass (/bæs/5), bus (/bʌs/), and boss (/bɑs/). The stimulus block was again repeated four times, for a total of 256 trials for the session.
Participants were tested individually in an anechoic chamber, seated in front of a MacBook Pro laptop running the experiment in the Praat ExperimentMFC format (Boersma and Weenink,
2.4 Analysis
In order to quantify the participants' use of the acoustic cues, a logistic regression analysis was performed on the obtained response data per session and per response category, using the glm() function in R (R Core Team, 2023). The model structure is as follows:
where P is the probability that a given response category (e.g., JP /aa/) is chosen, and 1 − P is the probability that the other three categories (e.g., JP /ee/, /e/, or /aa/) are chosen. The odds is log-transformed to fit a sigmoidal curve to the data, which is more appropriate than the straight line of a linear regression model for analyzing speech perception data. The intercept α is the bias coefficient, which reflects how likely the particular response category is to be chosen in general. The stimulus-tuned coefficients βs represent the extent to which the F1, F2, and duration steps, coded from “1” (smallest) to “4” (largest), cause a change in the likelihood of the response category being chosen.
2.5 Results
Table 1 shows the results of the logistic regression analyses on all participants' pooled data. The coefficients βF1, βF2, and βdur can be plotted to graphically represent the estimated locations of response categories in the stimulus space (Morrison,
Table 1
| Session | Vowel | α | βF1 | βF2 | βdur |
|---|---|---|---|---|---|
| JP | /ee/ | −6.943 | −0.261 | 0.988 | 0.667 |
| JP | /e/ | −2.167 | −0.309 | 0.749 | −0.570 |
| JP | /aa/ | −2.228 | 0.226 | −0.390 | 0.916 |
| JP | /a/ | 1.552 | 0.117 | −0.184 | −0.840 |
| EN | /ɛ/ | −5.293 | −0.386 | 1.404 | −0.108 |
| EN | /æ/ | −3.535 | 0.391 | 0.298 | 0.161 |
| EN | /ʌ/ | −1.313 | 0.158 | 0.150 | −0.195 |
| EN | /ɑ/ | 1.735 | −0.335 | −0.908 | 0.216 |
Results of logistic regression analyses on all participants' data in the experiment.
Figure 3

Plot of logistic regression coefficients in Table 1 (black = JP, white = EN).
Let us briefly examine the overall response patterns in the figure. Regarding the JP responses, the relative positions of /e/ and /a/ on the βF1-βF2 plane are as expected, since mid front /e/ should show lower βF1 and higher βF2 than low central /a/. Phonologically long /ee/ and /aa/ are proximal to their short counterparts in βF1 and βF2, but larger in βdur. This is consistent with the traditional description of Japanese long vowels as a sequence of two identical vowels at the phonological level. As for the EN responses, the relative positions of /ɛ/ and /ʌ/ are similar to those of JP /e/ and /a/, while /æ/ seems to be somewhat distant, on the βF1-βF2 plane. Far away from all other categories is /ɑ/, with very low βF1 and βF2. As for βdur, the four EN categories seem to occupy an intermediate position between JP long and short categories.
To further investigate the response patterns, linear mixed-effects (LME) models were applied to the by-participant results of the logistic regression analyses, using the lme4 (Bates et al.,
The model tests whether the response categories differ on a stimulus-tuned coefficient (βF1, βF2, or βdur) at a statistically significant level, controlling for the potential variability across participants and session order. Note that both JP and EN categories are included in the model, as the coefficients can in principle be compared across sessions, since the JP and EN stimuli share the same acoustic properties.
The LME model for βF1 with JP /a/ as the reference level showed significantly smaller estimates for EN /ɛ/ (β = −1.110, s.d. = 0.265, t = −4.180, p < 0.001) and /ɑ/ (β = −0.561, s.d. = 0.265, t = −2.116, p = 0.035), suggesting that the two EN categories are higher in perceived vowel height than the reference. The model for βF2 also yielded significantly larger estimates for EN /ɛ/ (β = 5.391, s.d. = 0.372, t = 14.492, p < 0.001) and /æ/ (β = 1.248, s.d. = 0.372, t = 3.355, p < 0.001), as well as a significantly smaller estimate for EN /ɑ/ (β = −0.892, s.d. = 0.372, t = −2.397, p = 0.017), than the reference JP /a/. This suggests that EN /ɛ/ and /æ/ are perceptually represented as more fronted, and EN /ɑ/ as more back, than JP /a/. No significant difference was found between EN /ʌ/ and JP /a/ in either βF1 or βF2. As for βdur, all EN categories had significantly larger estimates than the reference JP /a/ (p < 0.05 for EN /ʌ/ and ps < 0.001 for /ɛ/, /æ/, and /ɑ/). An additional LME model with EN /ʌ/ as reference found significantly larger βdur estimates for JP /ee/ (β = 1.646, s.d. = 0.378, t = 4.359, p < 0.001) and /aa/ (β = 1.423, s.d. = 0.378, t = 3.768, p < 0.001), but no significant difference was found for the other three EN categories. The results suggest that the four EN categories are represented with an intermediate perceptual duration between the long and short JP categories, with no significant difference between the EN categories themselves.
2.6 Interpretation
The above results can be interpreted as follows. First, AmE /ɛ/ and /ʌ/ are qualitatively assimilated to Japanese /e/ and /a/, given the similar βF1 and βF2 estimates between EN /ɛ/ and JP /e/ and between EN /ʌ/ and JP /a/, respectively. If a separate category had been formed for AmE /ɛ/, which is lower in phonetic height than Japanese /e/, then βF1 for EN /ɛ/ should have been larger than that for JP /e/, but this was not the case. Also, given the non-significant differences in βF1 and βF2 between EN /ʌ/ and JP /a/, it is unlikely that AmE /ʌ/ was reliably discriminated from Japanese /a/. In contrast, AmE /æ/ was most likely perceived as a separate category. Given its significantly larger βF2 than JP /a/, the AmE vowel may be represented as “a fronted version of /a/.” While these results are consistent with previous findings, it has additionally been shown that AmE /æ/ is distinguished from Japanese /a/ by the F2 cue and not by the F1 cue.
The result for EN /ɑ/, however, was somewhat unexpected. Although AmE /ɑ/ is reported to be qualitatively assimilated to Japanese /a/ (Strange et al., 1998), the βF1 and βF2 estimates for EN /ɑ/ responses were significantly lower than for JP /a/. There are a few possible explanations for this finding. First, the learners may have associated AmE /ɑ/ with Japanese /o/ at the orthographic level, since the AmE sound is often written with “o” (e.g., boss, lot, not), as is the Japanese sound when written in the Roman alphabet (e.g., bosu /bosu/ “boss”). This possibility is particularly plausible because the participants had learned English mostly in written rather than oral form. Second, the participants may have been referring to AmE /ɔ/ rather than /ɑ/ when they chose boss as their response. The experimental design assumed that the vowel in boss is /ɑ/ because of the widespread and ongoing low back merger in many dialects of AmE (Labov et al.,
Finally, it is worth noting that the duration cue was not utilized very actively in the EN session. Judging from their intermediate βdur between JP long and short categories, the AmE vowel categories appear to be unspecified in terms of phonological length. This result is consistent with Strange et al. (1998)'s finding that Japanese listeners did not show duration-based categorization when AmE vowels were presented in isolation as in the current experiment.
3 Simulation
Following the above experimental results, we now present in this section a formal computational modeling of how L1 Japanese listeners may develop a new sound category for L2 AmE /æ/ (or not for other categories) within the L2LP framework. We compare two versions of simulations using StOT, one segment- and the other feature-based, as they make divergent predictions about how L1 and L2 linguistic experience shapes listeners' perception. These predictions are compared with the experimental result to evaluate which version is more plausible. We begin by outlining the general procedure of the simulations, followed by the segmental and then by featural simulations.
3.1 General procedure
With StOT, speech perception can be modeled with a set of Optimality Theoretic, negatively formulated cue constraints (Escudero,
The ranking values of the constraints are not determined manually, but are learned computationally from the input data through the Gradual Learning Algorithm (GLA), an error-driven algorithm for learning optimal constraint rankings in StOT (Boersma and Hayes,
The segmental and featural versions of the simulations use the above two computational tools, with the same parameter settings whenever possible. All constraints have an initial ranking value of 100.0, and the evaluation noise is fixed at 2.0. The plasticity is initially set to 1.0, decreasing by a factor of 0.7 per virtual year. The number of yearly input tokens was 10,000. These settings are mostly taken from previous studies, Boersma and Escudero (
The input data for training the virtual listeners are randomly generated using the parameters in Table 2. The mean formant values are taken from Nishi et al. (
Table 2
| Language | Vowel | F1 (mel) | F2 (mel) | Frequency (%) | ||
|---|---|---|---|---|---|---|
| Mean | S.d. | Mean | S.d. | |||
| Japanese | /e/ | 573 | 100 | 1,421 | 150 | 33.3 |
| Japanese | /a/ | 758 | 100 | 1,086 | 150 | 33.3 |
| Japanese | /o/ | 533 | 100 | 841 | 150 | 33.3 |
| AmE | /ɛ/ | 721 | 50 | 1,368 | 100 | 25.0 |
| AmE | /æ/ | 792 | 50 | 1,363 | 100 | 25.0 |
| AmE | /ʌ/ | 724 | 50 | 1,144 | 100 | 25.0 |
| AmE | /ɑ/ ([ɑ]) | 824 | 50 | 1,145 | 100 | 12.5 |
| AmE | /ɑ/ ([ɔ]) | 749 | 50 | 1,037 | 100 | 12.5 |
Input training parameters for the simulations.
In the following two sections, we present how segmental and featural versions of virtual StOT listeners, trained with the same L1 Japanese and L2 AmE input, may develop a new category for /æ/ (and not for other AmE vowels), like the real listeners in our experiment. Each section begins with a brief illustration of cue constraints, namely cue-to-segment or cue-to-feature constraints. In line with the Full Transfer hypothesis of L2 acquisition (Schwartz and Sprouse, 1996), L2LP assumes that the initial state of L2 perception is a Full Copy of the end-state L1 grammar. Thus, we first train the perception grammar with Japanese input tokens for a total of 12 virtual years, which is copied to serve as the basis for L2 speech perception. Based on L2LP's further assumption that L2 learners have Full Access (Schwartz and Sprouse, 1996) to L1-like learning mechanisms, the copied perception grammar is then trained with AmE input in the same way, but with a decreased plasticity (1.0 × 0.712 = 0.014 at age 12, which further decreases by a factor of 0.7 per year).
3.2 Segmental simulation
3.2.1 Cue-to-segment constraints
Most previous studies aimed at formally modeling the process of L2 speech perception within L2LP (e.g., Escudero and Boersma,
Table 3
| [F1 = 850, F2 = 1,400] | [F2 = 1,400] */a/ | [F1 = 850] */e/ | [F1 = 850] */a/ | [F2 = 1,400] */e/ |
|---|---|---|---|---|
/e/ | * | * | ||
| /a/ | *! | * |
Example of segmental perception grammar.
Note, however, that the same vowel token will not always be perceived as /e/ due to the probabilistic nature of StOT. It is possible that in some cases the constraint “[F1 = 850 mel] */e/” will outrank “[F2 = 1,400 mel] */a/,” making /a/ as the alternative winner. The probability of such an evaluation is increased by GLA if and when the listener notices that the intended from should be /a/ rather than /e/ through their lexical knowledge and the semantic context (e.g., aki “autumn” should have been perceived instead of eki “station” given the conversational context). Table 4 illustrates how such learning takes place. Here, the ranking values of the constraints that led to the perception of the incorrect winner (“✓”) are increased (“←”), while the ranking values of the constraints that would lead to the correct form (“
”) are decreased (“ → ”), by the current plasticity value. This makes it more likely that the same token will be perceived as /a/ rather than /e/ in future evaluations.
Table 4
| [F1 = 850, F2 = 1,400] | [F2 = 1,400] */a/ | [F1 = 850] */e/ | [F1 = 850] */a/ | [F2 = 1,400] */e/ |
|---|---|---|---|---|
/e/ | ←* | ←* | ||
| ✓ /a/ | *! → | * → |
Constraint updating in segmental grammar.
3.2.2 L1 perception
Our virtual segmental learner starts with a “blank” perception grammar, which has a total of 96 cue-to-segment constraints (16 F1 bins + 16 F2 bins, multiplied by three segmental categories /e/, /a/, and /o/), all ranked at the same initial value of 100.0.7 The learner then begins to receive L1 input, namely random tokens of Japanese /e/, /a/, and /o/, which occur with equal frequency. The formant values of each vowel token is randomly determined based on the means and standard deviations in Table 2, which are then rounded to the nearest bins to be evaluated by the corresponding constraints. Whenever there is a mismatch between the perceived and intended forms, GLA updates the ranking values of the relevant cue constraints by adding or subtracting the current plasticity value.
Figure 4 shows the result of L1 learning. The grammar was tested 100 times on each combination of F1 and F2 bins. The vertical axis in the figure shows the probability of segmental categories being perceived given the F1-F2 bin combination, as calculated by logistic regression analyses as in (1) but without the duration coefficients. It can be seen that the virtual listener perceives /e/ when F1 is low and F2 is high, and /a/ when F1 is high and F2 is low, similar to the perception patterns of the real listeners in the experiment (cf. Figure 3). Note that /o/ can also be perceived when F1 and F2 are both very low.
Figure 4

Simulated segmental perception after learning Japanese as L1 for 12 years, with no representation for /æ/.
3.2.3 L2 perception
The segmental learner is then exposed to L2 AmE data for the first time in life. Following L2LP's Full Copying hypothesis, the 96 cue constraints and their ranking values are copied over. Given the experimental results, the L1 vowel labels /e/, /a/, and /o/ in the copied constraints are relabeled as L2 /ɛ/, /ʌ/, and /ɑ(ɔ)/, respectively. This alone would be sufficient to explain real learners' perception of seemingly L1-assimilated vowels: /ɛ/ (= /e/) is perceived when F1 is low and and F2 is high, /ʌ/ (= /a/) when F1 is high and F2 is low, and /ɑ(ɔ)/ (= /o/) when F1 and F2 are both very low (cf. Figures 3, 4).
There is a problem, however, in that the perception of AmE /æ/ cannot be adequately modeled by mere copying. Since the grammar can only perceive three existing segmental categories, a new category for /æ/ must be manually added to the grammar. The act of adding a new category itself is not theoretically unsupported, since learners may notice a lexical distinction denoted by the vowel contrast (e.g., bass vs. bus)—perhaps due to repeated communicative errors—which motivates them to form a new phonological category. However, we encounter a puzzle here: how would the lexical distinction between bass and bus help the listeners notice the phonological contrast between /æ/ and /ʌ/ if these words sound the “same” to them? Wouldn't these words simply be represented as homophones? For example, we can see in Figure 4 that most tokens of both the bass vowel /æ/, which typically has high F1 and F2, and the bus vowel /ʌ/, which typically has low F1 and F2, are perceived as Japanese /a/. Thus, leaving open the possibility of L2 listeners manually adding a new category still begs the question of what the precise mechanism that allows the learner to do so is.
Even if we ignore this theoretical problem and add 32 new constraints (16 F1 bins + 16 F2 bins) for /æ/ (e.g., “[F1 = 1400 mel] */æ/”) to the L2 grammar, we encounter another difficulty: The simulated learning outcome does not resemble actual perceptual behavior. In fact, the model overperforms. This can be seen in Figure 5, which shows the result after learning L2 AmE for six years. Despite the decreased plasticity, the grammar has learned to correctly perceive not only /æ/ but also other vowels /ɛ/, /ʌ/, and /ɑ/, according to the acoustic distributions of the input. This is clearly different from the real learners' perception observed in the experiment, where the latter three vowels /ɛ/, /ʌ/, and /ɑ/ were perceived as Japanese /e/, /a/, and /o/, respectively. The simulated learner therefore becomes too nativelike, showing almost identical perception patterns to those of an age-matched virtual L1 AmE listener (Figure 6). This is rather unrealistic, since very few adult L2 learners, let alone those at an intermediate level, are expected to exhibit nativelike perceptual performance.
Figure 5

Simulated segmental perception after subsequently learning AmE as L2 for 6 years.
Figure 6

Simulated segmental perception after learning only AmE as L1 for 18 years.
3.3 Featural simulation
3.3.1 Cue-to-feature constraints
Our featural simulation is based on Boersma and Chládková (
Table 5 shows how a featural Japanese grammar perceives a vowel token with [F1 = 850 mel] and [F2 = 1,400 mel] (i.e., an [æ]-like token) through two height features (/mid/ and /low/) and two backness features (/front/ and /central/). The candidates are four logical combinations of these features, two of which are well-formed in the L1 (/mid, front/ = /e/ and /low, central/ = /a/) and the other two of which are ill-formed (/mid, central/ and /low, front/). Structural constraints against ill-formed perceptual output are usually learned to be ranked very high, as is the case in the table, thus excluding the perception of /mid, central/ and /low, front/. The cue constraint “[F2 = 1,400 mel] */central/” then outranks “[F1 = 850 mel] */mid/,” making /mid, front/ the winner.
Table 5
| [F1 = 850, F2 = 1,400] | */mid, central/ | */low, front/ | [F2 = 1,400] */central/ | [F1 = 850] */mid/ | [F1 = 850] */low/ | [F2 = 1,400] */front/ | */low, central/ | */mid, front/ |
|---|---|---|---|---|---|---|---|---|
| /mid, front/ | * | * | * | |||||
| /mid, central/ | *! | * | * | |||||
| /low, front/ | *! | * | * | |||||
| /low, central/ | *! | * | * |
Example of featural perception grammar.
Perceptual learning in the featural grammar works in the same way as in the segmental grammar, as shown in Table 6. When the listener detects a mismatch between the intended form (“
”) and the perceived form (“✓”), GLA updates the grammar by increasing the ranking values of all constraints that led to the incorrect winner (“←”) and decreasing the ranking values of the constraints that would lead to the correct form (“ → ”) by the current plasticity. Note that both cue and structural constraints are learned.
Table 6
| [F1 = 850, F2 = 1,400] | */mid, central/ | */low, front/ | [F2 = 1,400] */central/ | [F1 = 850] */mid/ | [F1 = 850] */low/ | [F2 = 1,400] */front/ | */low, central/ | */mid, front/ |
|---|---|---|---|---|---|---|---|---|
/mid, front/ | ←* | ←* | ←* | |||||
| ✓ /low, central/ | *! → | * → | * → |
Constraint updating in featural grammar.
3.3.2 L1 perception
Just like the segmental learner, our featural learner starts with a “blank” perception grammar, which has 80 cue constraints (16 F1-to-height constraints for each of two height features /mid/ and /low/, and 16 F2-to-backness constraints for each of 3 backness features /front/, /central/, and /back/) as well as 6 structural constraints (two height features × three backness features), all ranked at the same initial value of 100.0.9 The learner then begins to receive L1 input, namely randomly generated tokens of Japanese /e/ (/mid, front/), /a/ (/low, central/), and /o/ (/mid, back/), as per Table 2. The correspondence between features and categories (e.g., /a/ = /low, central/) is based on Boersma and Chládková (
Figure 7 shows the result of L1 learning, tested in the same way as the segmental grammar. A notable difference from the segmental result (Figure 4) is that the featural grammar can perceive a feature combination that does not occur in the L1 input, despite the high-ranked structural constraints against such ill-formed output. For example, a token with high F1 and F2, which the segmental grammar perceived as /a/ most of the time or as /e/ otherwise, can sometimes be perceived as /low, front/, which has no segmental equivalent in Japanese. What this means is that the featural grammar may prefer to perceive a structurally ill-formed form such as /low, front/ over well-formed forms such as /low, central/ if there is sufficient cue evidence to support the evaluation. This essentially expresses the perceptual deviance of [æ] that segmental modeling fails to capture: The vowel is too /front/ to be /low, central/ (= /a/).
Figure 7

Simulated featural perception after learning Japanese as L1 for 12 years.
3.3.3 L2 perception
The featural learner then begins to learn L2 AmE. Since the initial L2 grammar is a copy of the L1 grammar, it has 80 cue constraints and 6 structural constraints with the copied ranking values. Following the experimental results, and to make the featural simulation compatible with the segmental one, we assume that L2 /ɛ/ is represented as /mid, front/, /ʌ/ as /low, central/, and /ɑ(ɔ)/ as /mid, back/ in the grammar. No additional constraint is needed to model /æ/ (/low, front/).
Figure 8 shows the result of learning L2 AmE for six years. It can be seen that the feature combination /low, front/ is much more likely to be perceived than it was in Figure 7 because the ranking value of the structural constraint “*/low, front/” has decreased. The weakening of the constraint occurred because in the L2 AmE environment, the features /low/ and /front/ often co-occur, and /low, front/ should be lexically distinguished from other feature combinations for successful communication. A new category therefore emerged from existing features by improving the well-formedness of the once ill-formed feature combination, without resorting to any L2-specific manipulation of the grammar as in the segmental modeling.
Figure 8

Simulated featural perception after subsequently learning AmE as L2 for 6 years.
Another notable finding is that the simulated perception in Figure 8 differs from the simulated L1 AmE perception in Figure 9. One salient difference lies in /ʌ/, which was learned as /low, central/ by the learner grammar, whereas it is represented as /mid, back/ in the native grammar.10 The native perception is symmetrical because it reflects the production environment of AmE vowels, whereas the learner perception is asymmetrical because L2 AmE sounds are perceived through copied L1 Japanese features. The simulated learner perception actually resembles the real learners' perception, especially in the use of the F2 cue, where /ɛ/ (/mid, front/) and /æ/ (/low, front/) are perceptually more fronted, while /ɑ/ (/mid, back/) is more back, than /ʌ/ (/low, central/).
Figure 9

Simulated featural perception after learning only AmE as L1 for 18 years.
4 General discussion
This study examined how L1 Japanese learners of L2 AmE develop a new phonological representation for /æ/ by employing experimental and computational-phonological approaches. The experimental results suggested that AmE /æ/ is represented as a separate category by intermediate-level learners, distinguished from Japanese /a/ based on the F2 cue, while adjacent AmE /ɛ/, /ʌ/, and /ɑ/ are assimilated to Japanese /e/, /a/, and /o/, respectively. To explain and replicate these results with the L2LP model, segment- and feature-based versions of perceptual simulations were performed using StOT and GLA. The segmental modeling was theoretically inadequate because it failed to elucidate the mechanism for noticing the perceptual distinctness of /æ/, and was also practically implausible because the predicted overall perception patterns were too native AmE-like compared to real learners' perception. In contrast, the featural modeling explained the emergence of a new category for AmE /æ/ and the lack thereof for /ɛ/, /ʌ/, and /ɑ/ by assuming that L2 sounds are perceived through copied L1 features, i.e., */low, front/ vs. /mid, front/, /low, central/, and /mid, back/, respectively. The simulated learning outcome closely resembled real perception.
In this section, we discuss the implications of the above findings for L2LP and the other two dominant models of L2 perception, as well as directions for extending the current study in future research.
4.1 Implications for L2 perception models
4.1.1 L2LP
Our simulations have shown that, similar to the Unfamiliar New scenario in Escudero and Boersma (
One important point about the feature-based modeling is that the relationship between acoustic cues and phonological features is considered to be language-specific and relative. For example, while the F1 cue may be mapped to three height features (/high/, /mid/, and /low) in many languages, in some languages such as Portuguese and Italian there are four target heights (/high/, /mid-high/, /mid-low/, and /low/) and in others such as Arabic and Quechua there are only two (/high/ and /low/). Also, even if two languages share the “same” set of height features, what is perceived as /high/ in one language may not be also perceived as also /high/ in another, since the actual F1 values of high vowels varies across languages or even language varieties (Chládková and Escudero,
4.1.2 SLM
While the current study aimed to explain the process of new category formation within the framework of L2LP, the results also have useful implications for SLM. Specifically, it can be proposed that cross-linguistic categorical dissimilarity is defined as a mismatch of existing L1 features (e.g., */low, front/), with the caveat that the actual phonetic property of a feature is language-specific and relative as discussed above. This proposal is actually compatible with one of the hypotheses (H6) of the original SLM (Flege,
Much of this discussion, however, did not find its way into the revised SLM (Flege and Bohn,
4.1.3 PAM
The implication of feature-based modeling for PAM is similar to that for SLM: Cross-linguistic dissimilarity can be defined as featural discrepancy. However, the implication is unique for PAM because, unlike SLM and L2LP which model speech perception as the abstraction of acoustic cues into sound representations (be them segments or features), PAM subscribes to a direct realist view that listeners directly perceive articulatory gestures of the speaker. PAM also distinguishes between phonetic and phonological levels of representations (like L2LP, in a broad sense), whereas SLM defines sound categories strictly at the phonetic or allophonic level. These differences in theoretical assumption raise a crucial question about what features really are: Are they articulatory or auditory, and phonetic or phonological? As mentioned earlier, the current study assumed what Boersma and Chládková (
4.2 Future directions
Having discussed the theoretical implications of the feature-based approach, we now address how the current modeling can be practically extended to improve its adequacy in future research. First, as for the acoustic cues, we chose not to include duration because the participants in our experiment do not seem to have used it to categorize the target L2 AmE vowels, but it remains to be modeled why L1 listeners of Japanese with phonological vowel length would show such perception patterns. This can actually be a task effect, since duration-based categorization of AmE vowels into Japanese long and short ones was only observed when AmE vowels were embedded within a carrier sentence (Strange et al., 1998), i.e., when the target vowel duration could be compared to the duration of other vowels in the carrier sentence (cf. within-language feature relativity in Section 4.1.1). Thus, the modeling may need to incorporate some kind of temporal normalization to explain the potential task dependency. Escudero and Bion (
We also believe that further empirical testing is needed to complement the formal computational modeling. One limitation of the current experiment, or behavioral experiments in general, is that features cannot be directly observed in participant responses. To overcome this weakness, neural studies as in Scharinger et al. (2011) or Mesgarani et al. (
5 Conclusion
This study proposed that perceived (dis)similarity between L1 and L2 sounds, which is considered crucial for the process of new L2 category formation but has long remained elusive, can be better defined by assuming feature-level representations as the fundamental unit of perception, rather than segmental categories. Through our formal modeling based on L2LP and StOT, we argued that an L2 sound (e.g., AmE /æ/) whose Familiar acoustic cues (e.g., F1 and F2) map to a bundle of L1 features that is structurally ill-formed (e.g., */low, front/ in Japanese) is perceived as deviant and thus is subject to category formation, whereas an L2 sound (e.g., AmE /ɛ/, /ʌ/, and /ɑ/) whose cues map to a well-formed L1 feature bundle (e.g., /mid, front/, /low, central/, and /mid, back/ in Japanese) is prone to assimilation, regardless of the ostensible acoustic distance between L1 and L2 segmental categories. The proposed feature-based modeling was consistent with our experimental results, where real L1 Japanese listeners seem to have established a distinct representation for L2 AmE /æ/ but not for /ɛ/, /ʌ/, and /ɑ/, which the segment-based modeling failed to predict and replicate. While feature-based approaches to L2 learning are still scarce compared to the vast literature on segment-based approaches, perhaps because the intangible nature of features cannot be captured without a computational platform that is currently only available to L2LP, the benefits of adopting and extending the approach are expected to go beyond the model (e.g., SLM and PAM) and beyond the current learning scenario (i.e., other sound contrasts in different language combinations), the pursuit of which should ultimately help deepen our understanding of how L2 speech acquisition proceeds as a whole.
Statements
Data availability statement
The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.
Ethics statement
The studies involving humans were approved by Academic Research Ethical Review Committee, Waseda University. The studies were conducted in accordance with the local legislation and institutional requirements. The participants provided their written informed consent to participate in this study.
Author contributions
KY: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Writing – original draft, Writing – review & editing. JW: Conceptualization, Formal analysis, Funding acquisition, Investigation, Methodology, Software, Validation, Writing – original – draft. MK: Conceptualization, Funding acquisition, Project administration, Supervision, Writing – original draft. PE: Conceptualization, Formal analysis, Funding acquisition, Project administration, Supervision, Validation, Writing – original draft.
Funding
The authors declare financial support was received for the research, authorship, and/or publication of this article. KY and MK's work was funded by JSPS Grant-in-Aid for Scientific Research (grant number: 21H00533). JW's work was funded by Samsung Electronics Co., Ltd. (grant number: A0342-20220008); the funder was not involved in the study design, collection, analysis, interpretation of data, the writing of this article, or the decision to submit it for publication. PE's work was funded by ARC Future Fellowship grant (grant number: FT160100514).
Acknowledgments
The authors thank Mikey Elmers for volunteering to provide his voice for the AmE stimuli.
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Footnotes
1.^Our view of features departs from the generally assumed universal set of binary phonological features. Specifically, we assume in this paper that features are language-specific and emergent rather than universal and innate (Boersma et al.,
2.^AmE is considered as the target variety of English because it is widely used in the formal English language education in Japan and is most familiar to the learners (Sugimoto and Uchida, 2020).
3.^Japanese has five vowel qualities /i/, /e/, /a/, /o/, and /u/, which form five short (1-mora) and long (2-mora) pairs. Long vowels are transcribed with double letters (e.g., /aa/) in this study because they can underlyingly be a sequence of two identical vowels. The transcription “/a(a)/” here indicates “either /a/ or /aa/” because the AmE vowel was perceived as Japanese /a/ when presented in isolation but as Japanese /aa/ when embedded in a carrier sentence.
4.^Similar to AmE /æ/, AmE /ɑ/ was perceived as Japanese /a/ when presented in isolation but as Japanese /aa/ when embedded in a carrier sentence.
5.^Participants were reminded that the pronunciation of bass was not /beɪs/ “low frequency sound” but /bæs/ “a type of fish” in a short practice before the EN session, where natural tokens of the four English words were used as tokens. The JP session also followed a practice with natural tokens of the four Japanese words as tokens. The Japanese and English tokens were produced by the same speakers as those in Section 2.2.
6.^The distinction between lexical and semantic levels of representations goes beyond the scope of our simulations; see Boersma (
7.^The grammar is not truly “blank” because it already knows three segmental categories onto which the cues are mapped. Boersma et al. (
8.^Constraints that map acoustic cues to privative features were first introduced by Boersma et al. (
9.^This grammar is also not truly “blank” because it already knows two height and three backness features. Boersma et al. (
10.^We used the same set of feature labels in both native and learner grammars to allow for a direct comparison between them, not because we assume a universal set of features across all languages (cf. Section 4.1.1). For example, we could relabel the /mid/ feature in the native grammar as /low-mid/ and still get the same result as Figure 9.
11.^Traditional OT grammars can be seen as a special case of StOT grammars, with integer ranking values and zero evaluation noise.
References
1
ArchibaldJ. (2023). Differential substitution: a contrastive hierarchy account. Front. Lang. Sci. 2, 1–13. 10.3389/flang.2023.1242905
2
BalasA. (2018). Non-Native Vowel Perception: The Interplay of Categories and Features. Poznań: Wydawnictwo Naukowe UAM.
3
BatesD.MächlerM.BolkerB. M.WalkerS. C. (2015). Fitting linear mixed-effects models using lme4. J. Stat. Softw. 67, 1–48. 10.18637/jss.v067.i01
4
BendersT.EscuderoP.SjerpsM. J. (2012). The interrelation between acoustic context effects and available response categories in speech sound categorization. J. Acoust. Soc. Am. 131, 3079–3087. 10.1121/1.3688512
5
BestC. T. (1995). “A direct realist view of cross-language speech perception,” in Speech Perception and Linguistic Experience: Issues in Cross-Language Research, ed StrangeW. (Timonium, MD: York Press), 171–204.
6
BestC. T.TylerM. (2007). “Nonnative and second-language speech perception: commonalities and complementarities,” in Language Experience in Second Language Speech Learning: In Honor of James Emil Flege, ed BohnO. S.MunroM. J. (Amsterdam; Philadelphia, PA: John Benjamins), 13–34.
7
BoersmaP. (1998). Functional Phonology: Formalizing the Interactions between Articulatory and Perceptual Drives (Ph.D. thesis), University of Amsterdam. The Hague: Holland Academic Graphics.
8
BoersmaP. (2009). “Cue constraints and their interactions in phonological perception and production,” in Phonology in Perception, eds BoersmaP.HamannS. (Berlin: De Gruyter), 55–110.
9
BoersmaP. (2011). “A programme for bidirectional phonology and phonetics and their acquisition and evolution,” in Bidirectional Optimality Theory, eds BenzA.MattauschJ. (Amsterdam: John Benjamins), 33–72.
10
BoersmaP.ChládkováK. (2011). “Asymmetries between speech perception and production reveal phonological structure,” in Proceedings of the 17th International Congress of Phonetic Sciences, eds LeeW.-S.ZeeE. (Hong Kong: The University of Hong Kong), 328–331.
11
BoersmaP.ChládkováK.BendersT. (2022). Phonological features emerge substance-freely from the phonetics and the morphology. Can. J. Linguist. 67, 611–669. 10.1017/cnj.2022.39
12
BoersmaP.EscuderoP. (2008). “Learning to perceive a smaller L2 vowel inventory: an Optimality Theory account,” in Contrast in Phonology: Theory, Perception, Acquisition, eds P. Avery, E. Dresher, and K. Rice (Berlin: de Gruyter), 271–302.
13
BoersmaP.EscuderoP.HayesR. (2003). “Learning abstract phonological from auditory phonetic categories: an integrated model for the acquisition of language-specific sound categories,” in Proceedings of the 15th International Congress of Phonetic Sciences, eds M. J. Solé, D. Recasens, and J. Romero (Barcelona), 1013–1016.
14
BoersmaP.HayesB. (2001). Empirical tests of the gradual learning algorithm. Linguist. Inq. 32, 45–86. 10.1162/002438901554586
15
BoersmaP.WeeninkD. (2023). Praat: Doing Phonetics by Computer. Amsterdam: University of Amsterdam.
16
ChenJ.ChangH. (2022). Sketching the landscape of speech perception research (2000–2020): a bibliometric study. Front. Psychol. 13, 1–14. 10.3389/fpsyg.2022.822241
17
ChládkováK.BoersmaP.BendersT. (2015a). “The perceptual basis of the feature vowel height,” in Proceedings of the 18th International Congress of Phonetic Sciences, ed The Scottish Consortium for ICPhS 2015 (Glasgow: The University of Glasgow), 711.
18
ChládkováK.BoersmaP.EscuderoP. (2022). Unattended distributional training can shift phoneme boundaries. Bilingual. Lang. Cognit. 25, 827–840. 10.1017/S1366728922000086
19
ChládkováK.EscuderoP. (2012). Comparing vowel perception and production in Spanish and Portuguese: European versus Latin American dialects. J. Acoust. Soc. Am. 131, EL119–EL125. 10.1121/1.3674991
20
ChládkováK.EscuderoP.LipskiS. C. (2015b). When “AA” is long but “A” is not short: speakers who distinguish short and long vowels in production do not necessarily encode a short long contrast in their phonological lexicon. Front. Psychol. 6, 1–8. 10.3389/fpsyg.2015.00438
21
EscuderoP. (2005). Linguistic Perception and Second Language Acquisition: Explaining the Attainment of Optimal Phonological Categorization (PhD thesis). Utercht University. Amsterdam: LOT.
22
EscuderoP. (2007). “Second-language phonology: the role of perception,” in Phonology in Context, ed PenningtonM. C. (London: Palgrave Macmillan), 109–134.
23
EscuderoP. (2009). “The linguistic perception of SIMILAR L2 sounds,” in Phonology in Perception, eds BoersmaP.HamannS. (Berlin: de Gruyter), 151–190.
24
EscuderoP. (2015). Orthography plays a limited role when learning the phonological forms of new words: the case of Spanish and English learners of novel Dutch words. Appl. Psycholinguist. 36, 7–22. 10.1017/S014271641400040X
25
EscuderoP.BionR. (2007). “Modeling vowel normalization and sound perception as sequential processes,” in Proceedings of the 16th International Congress of Phonetic Sciences, eds J. Trouvain, and W. J. Barry (Saarbrücken: Saarland University), 1413–1416.
26
EscuderoP.BoersmaP. (2004). Bridging the gap between L2 speech perception research and phonological theory. Stud. Sec. Lang. Acquisit. 26, 551–585. 10.1017/S0272263104040021
27
EscuderoP.KasteleinJ.WeiandK.van SonR. J. J. H. (2007). “Formal modelling of L1 and L2 perceptual learning: computational linguistics versus machine learning,” in Proceedings of the 8th Annual Conference of the International Speech Communication Association (Antwerp: International Speech Communication Association), 1889–1892.
28
EscuderoP.SimonE.MulakK. E. (2014). Learning words in a new language: orthography doesn't always help. Biling. Lang. Cogn. 17, 384–395. 10.1017/S1366728913000436
29
EscuderoP.WanrooijK. (2010). The effect of L1 orthography on non-native vowel perception. Lang. Speech53, 343–365. 10.1177/0023830910371447
30
Escudero P. Yazawa K. (in press). “The second language linguistic perception model (L2LP),” in The Cambridge Handbook of Bilingual Phonetics Phonology, ed M. Amengual (Cambridge: Cambridge University Press).
31
FlegeJ. E. (1995). “Second language speech learning: theory, findings, and problems,” in Speech Perception and Linguistic Experience: Issues in Cross-Language Research, ed StrangeW. (Timonium, MD: York Press), 233–277.
32
FlegeJ. E.BohnO.-S. (2021). “The revised speech learning model (SLM-r),” in Second Language Speech Learning: Theoretical and Empirical Progress, ed WaylandR. (Cambridge: Cambridge University Press), 3–83.
33
GreenbergS.ChristiansenT. U. (2019). The perceptual flow of phonetic information. Attent. Percept. Psychophys. 81, 884–896. 10.3758/s13414-019-01666-y
34
HamannS. (2009). “Variation in the perception of an L2 contrast: a combined phonetic and phonological account,” in Variation and Gradience in Phonetics and Phonology, eds F. Kügler, C. Féry, and R. van de Vijver (Berlin: De Gruyter), 71–98.
35
HamannS.ColomboI. E. (2017). A formal account of the interaction of orthography and perception. Nat. Lang. Linguist. Theory35, 683–714. 10.1007/s11049-017-9362-3
36
HirataY. (2004). Training native English speakers to perceive Japanese length contrasts in word versus sentence contexts. J. Acoust. Soc. Am. 116, 2384–2394. 10.1121/1.1783351
37
HirataY. (2017). “Second language learners' production of geminate consonants in Japanese,” in The Phonetics and Phonology of Geminate Consonants, ed KubozonoH. (Oxford: Oxford University Press), 163–184.
38
HisagiM.HigbyE.ZandonaM.KentJ.CastilloD.DavidovichI.et al. (2021). “Perceptual discrimination measure of non-native phoneme perception in early and late Spanish-English & Japanese-English bilinguals,” in Proceedings of Meetings on Acoustics, Vol. 42 (The Acoustical Society of America), 1–13.
39
IversonP.KuhlP.Akahane-YamadaR.DieschE.TohkuraY.KettermannA.et al. (2003). A perceptual interference account of acquisition difficulties for non-native phonemes. Cognition87, B47–B57. 10.1016/S0010-0277(02)00198-1
40
KawaharaH. (2006). STRAIGHT, exploitation of the other aspect of VOCODER: Perceptually isomorphic decomposition of speech sounds. Acoust. Sci. Technol. 27, 349–353. 10.1250/ast.27.349
41
KuhlP. K.ConboyB. T.Coffey-CorinaS.PaddenD.Rivera-GaxiolaM.NelsonT. (2008). Phonetic learning as a pathway to language: new data and native language magnet theory expanded (NLM-e). Philos. Transact. R. Soc. B Biol. Sci. 363, 979–1000. 10.1098/rstb.2007.2154
42
KuznetsovaA.BrockhoffP. B.ChristensenR. H. B. (2017). lmerTest package: tests in linear mixed effects models. J. Stat. Softw. 82, 1–26. 10.18637/jss.v082.i13
43
LabovW.AshS.BobergC. (2006). The Atlas of North American English: Phonetics, Phonology and Sound Change. Berlin: De Gruyter.
44
LambacherS. G.MartensW. L.KakehiK.MarasingheC. A.MolholtG. (2005). The effects of identification training on the identification and production of American English vowels by native speakers of Japanese. Appl. Psycholinguist. 26, 227–249. 10.1017/S0142716405050150
45
McAllisterR.FlegeJ. E.PiskeT. (2002). The influence of L1 on the acquisition of Swedish quantity by native speakers of Spanish, English and Estonian. J. Phon. 30, 229–258. 10.1006/jpho.2002.0174
46
MesgaraniN.CheungC.JohnsonK.ChangE. F. (2014). Phonetic feature encoding in human superior temporal gyrus. Science343, 1006–1010. 10.1126/science.1245994
47
MorrisonG. S. (2007). “Logistic regression modelling for first-and second-language perception data,” in Segmental and Prosodic Issues in Romance Phonology, eds M. J. Solé, P. Prieto, J. Mascaró, and M. J. Sol? (Amsterdam: John Benjamins), 219–236.
48
NishiK.StrangeW.Akahane-YamadaR.KuboR.Trent-BrownS. A. (2008). Acoustic and perceptual similarity of Japanese and American English vowels. J. Acoust. Soc. Am. 124, 576–588. 10.1121/1.2931949
49
PajakB.LevyR. (2014). The role of abstraction in non-native speech perception. J. Phon. 46, 147–160. 10.1016/j.wocn.2014.07.001
50
PrinceA.SmolenskyP. (1993). Optimality Theory: Constraint Interaction in Generative Grammar. Rutgers University Center for Cognitive Science Technical Report 2. New Jersey.
51
R Core Team (2023). R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing. Vienna.
52
SaitoK.van PoeterenK. (2018). The perception-production link revisited: the case of Japanese learners' English /ɹ/ performance. Int. J. Appl. Linguist. 28, 3–17. 10.1111/ijal.12175
53
ScharingerM.IdsardiW. J.PoeS. (2011). A comprehensive three-dimensional cortical map of vowel space. J. Cogn. Neurosci. 23, 3972–3982. 10.1162/jocn_a_00056
54
SchwartzB. D.SprouseR. A. (1996). L2 cognitive states and the full transfer/full access model. Sec. Lang. Res. 12, 40–72. 10.1177/026765839601200103
55
ShaferV. L.KreshS.ItoK.HisagiM.VidalN.HigbyE.et al. (2021). The neural timecourse of American English vowel discrimination by Japanese, Russian and Spanish second-language learners of English. Biling. Lang. Cogn. 24, 642–655. 10.1017/S1366728921000201
56
ShawJ. A.KawaharaS. (2017). The lingual articulation of devoiced /u/ in Tokyo Japanese. J. Phon. 66, 100–119. 10.1016/j.wocn.2017.09.007
57
ShinoharaY.HanC.HestvikA. (2019). “Effects of perceptual assimilation: the perception of English /æ/, /ʌ/, and /ɑ/ by Japanese speakers,” in Proceedings of the 19th International Congress of Phonetic Sciences, eds CalhounS.EscuderoP.TabainM.WarrenP. (Melbourne, VIC. Australasian Speech Science and Technology Association Inc.), 2344–2348.
58
ShinoharaY.HanC.HestvikA. (2022). Discriminability and prototypicality of nonnative vowels. Stud. Second Lang. Acquis.44, 1260–1278. 10.1017/S0272263121000978
59
ShinoharaY.IversonP. (2021). The effect of age on English /r/-/l/ perceptual training outcomes for Japanese speakers. J. Phon. 89, 1–24. 10.1016/j.wocn.2021.101108
60
StrangeW.Akahane-YamadaR.KuboR.Trent-BrownS. A.NishiK.JenkinsJ. J. (1998). Perceptual assimilation of American English vowels by Japanese listeners. J. Phon. 26, 311–344. 10.1006/jpho.1998.0078
61
SugimotoJ.UchidaY. (2020). English phonetics and teacher training: designing a phonetics course for Japanese preservice teachers. J. Phonet. Soc. Jpn24, 22–35. 10.24467/onseikenkyu.24.0_22
62
TsukadaK. (2012). Comparison of native versus nonnative perception of vowel length contrasts in Arabic and Japanese. Appl. Psycholinguist. 33, 501–516. 10.1017/S0142716411000452
63
TsukadaK.CoxF.HajekJ.HirataY. (2018). Non-native Japanese learners' perception of consonant length in Japanese and Italian. Sec. Lang. Res. 34, 179–200. 10.1177/0267658317719494
64
van LeussenJ.-W.EscuderoP. (2015). Learning to perceive and recognize a second language: The L2LP model revised. Front. Psychol. 6, 1–12. 10.3389/fpsyg.2015.01000
65
WhangJ.YazawaK. (2023). Modeling a phonotactic approach to segment recovery: the case of Japanese high vowels. Stud. Phonet. Phonol. Morphol. 29, 271–295. 10.17959/sppm.2023.29.2.271
66
YazawaK. (2020). Testing Second Language Linguistic Perception : A Case Study of Japanese, American English, and Australian English Vowels (Ph.D. thesis). Waseda University, Tokyo.
67
Yazawa K. (in press). NEW sounds can be easier to learn than SIMILAR sounds, but only when acoustic cues are FAMILIAR. Tsukuba Eng. Stud. 42.
68
YazawaK.WhangJ.EscuderoP. (2023). Australian English listeners' perception of Japanese vowel length reveals underlying phonological knowledge. Front. Psychol. 14, 1–14. 10.3389/fpsyg.2023.1122471
69
YazawaK.WhangJ.KondoM.EscuderoP. (2020). Language-dependent cue weighting: an investigation of perception modes in L2 learning. Sec. Lang. Res. 36, 557–581. 10.1177/0267658319832645
Summary
Keywords
Second Language Linguistic Perception (L2LP) model, Stochastic Optimality Theory, Gradual Learning Algorithm, category formation, features, computational modeling, Japanese, American English
Citation
Yazawa K, Whang J, Kondo M and Escudero P (2023) Feature-driven new sound category formation: computational implementation with the L2LP model and beyond. Front. Lang. Sci. 2:1303511. doi: 10.3389/flang.2023.1303511
Received
28 September 2023
Accepted
22 November 2023
Published
20 December 2023
Volume
2 - 2023
Edited by
Baris Kabak, Julius Maximilian University of Würzburg, Germany
Reviewed by
Mathias Scharinger, University of Marburg, Germany
Silke Hamann, University of Amsterdam, Netherlands
Updates

Check for updates
Copyright
© 2023 Yazawa, Whang, Kondo and Escudero.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Kakeru Yazawa yazawa.kakeru.gb@u.tsukuba.ac.jpJames Whang jamesw@snu.ac.kr
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.
/e/