Abstract
Speech-in-noise (SIN) perception is a complex cognitive skill that affects social, vocational, and educational activities. Poor SIN ability particularly affects young and elderly populations, yet varies considerably even among healthy young adults with normal hearing. Although SIN skills are known to be influenced by top-down processes that can selectively enhance lower-level sound representations, the complementary role of feed-forward mechanisms and their relationship to musical training is poorly understood. Using a paradigm that minimizes the main top-down factors that have been implicated in SIN performance such as working memory, we aimed to better understand how robust encoding of periodicity in the auditory system (as measured by the frequency-following response) contributes to SIN perception. Using magnetoencephalograpy, we found that the strength of encoding at the fundamental frequency in the brainstem, thalamus, and cortex is correlated with SIN accuracy. The amplitude of the slower cortical P2 wave was previously also shown to be related to SIN accuracy and FFR strength; we use MEG source localization to show that the P2 wave originates in a temporal region anterior to that of the cortical FFR. We also confirm that the observed enhancements were related to the extent and timing of musicianship. These results are consistent with the hypothesis that basic feed-forward sound encoding affects SIN perception by providing better information to later processing stages, and that modifying this process may be one mechanism through which musical training might enhance the auditory networks that subserve both musical and language functions.
Introduction
Understanding the neural bases of good speech-in-noise (SIN) perception during development, adulthood, and into old age is both clinically and scientifically important. However, it is challenging due to the complexity of the skill, which can be considered as a special case of auditory scene analysis, and can involve multiple cognitive processes depending on the information that is offered including spatial location, spectral and temporal regularity, and modulation (Moore and Gockel, ; Pressnitzer et al., ), and can be aided by visual cues (Suied et al., 2009), by predictions formed with the motor system (Du et al., ), and based on prior knowledge such as of language (Pickering and Garrod, ; Golestani et al., ) that can be used to constrain the interpretation of noisy information (Bendixen, ). The contribution of the fidelity with which an individual encodes various sound properties, including periodicity, which varies according to a variety of life experiences (e.g., musical experience Anderson et al., ) is not yet clear.
One means of observing the inter-individual differences in how people encode periodic characteristics of sound is the frequency-following response (FFR), an evoked response that is an index of the temporal representation of periodic sound in the brainstem (Chandrasekaran and Kraus, ; Skoe and Kraus, 2010), thalamus, and auditory cortex (Coffey et al., ,). Differences in the strength and fidelity of the fundamental frequency (f0) of the FFR have been linked to SIN perception such that increased FFR amplitude is associated with better performance (reviewed in Du et al., , see also Parbery-Clark et al., ; Anderson et al., ). However, enhancements and deficits of neural correlates that are related to SIN perception are most consistently observed either when the FFR is measured in very challenging listening conditions (e.g., Parbery-Clark et al., ), in the degree of degradation of the FFR signal between quiet and noisy conditions (e.g., Cunningham et al., ; Parbery-Clark et al., ; Song et al., 2011), or in the magnitude of enhancement that is conferred by predictability within a sound stream (e.g., Parbery-Clark et al., ). f0 representation in the FFR may be enhanced by training (Song et al., 2008, 2012) and is often observed to be stronger among musicians even to sounds presented in silence (e.g., Musacchia et al., ), suggesting that learning mechanisms related to identifying task-relevant features and possibly attention might act to bias and enhance incoming acoustic information and suppress noise (Suga, 2012).
However, a clear picture has not yet emerged (Coffey et al., ). It is unclear if the sometimes-observed relationship between the FFR and SIN is due only to better top-down mechanisms such as better stream segregation (Başkent and Gaudrain, ) or selective auditory attention (Parbery-Clark et al., ; Song et al., 2011; Lehmann and Schönwiesner, ), or if enhanced feed-forward stimulus encoding also plays a role. Here, we aimed to better understand the neural bases of periodicity coding in the brain under optimal listening conditions to understand its relevance to SIN; if basic encoding of sound quality in silence is important for more complex tasks such as hearing in noise, then we predict that there should be a relationship between FFR measured in silence and SIN performance. A secondary question we address is whether this relationship might be enhanced by musicianship.
Musicians are thought to have both enhanced bottom-up (Musacchia et al., ; Bidelman and Weiss, ) and top-down (Strait et al., 2010; Kraus et al., ) processing of sound. Because SIN perception and measures of basic sound encoding are related to musicianship, musical training has been proposed as a means of ameliorating poor SIN performance (reviewed in: Alain et al., ). Musical training places high demands on sensory, motor, and cognitive processing mechanisms that overlap between music and speech perception, and offers extensive repetition and emotional reward, which could stimulate auditory system enhancements that in turn impact speech processing (Patel, ). Several longitudinal studies support a causal relationship between musical training and SIN skills (Tierney et al., 2013; Kraus et al., ; Slater et al., 2015), although it is difficult to maintain full, experimental control over naturalistic training studies (Evans et al., ). A number of cross-sectional studies have also reported a musician advantage in SIN perception (Parbery-Clark et al., ,, , ; Strait et al., 2012; Zendel and Alain, 2012; Swaminathan et al., 2015); however, other studies have not found significant group differences (Ruggles et al., ; Boebinger et al., ) or have found the musicianship effect to be dependent upon the specific SIN task variations, such as the degree of information masking (Swaminathan et al., 2015; Başkent and Gaudrain, ) or the degree of reliance on pitch cues (Fuller et al., ). A recent review of SIN perception among musicians concluded that on balance there is good evidence for musician enhancement of SIN, but also highlighted the diversity of study designs used to study hearing in noisy conditions, which may contribute to inconsistent findings (Coffey et al., ). However, it remains uncertain to which aspects of cognition any musician advantage is owed: top-down processes such as selective attention and working memory that modulate early levels (Rinne et al., ) to filter and temporarily store incoming information (Strait and Kraus, 2011; Kraus et al., ), relatively immutable factors such as non-verbal IQ (Boebinger et al., ) that might affect multiple cognitive processes, or differences in basic sound encoding (reviewed in Anderson and Kraus, ; Du et al., ; Alain et al., , see also Weiss and Bidelman, 2015).
In the present study, we first aimed to clarify whether robust f0 encoding in the auditory system, which is known to be enhanced in musicians (Musacchia et al., ; Bidelman et al., ,), influences SIN perception in a feed-forward fashion. Rather than relating fundamental encoding recorded in the presence of noise to later performance (which has previously been shown, described above) and might include influences from top-down processes that spontaneously act to separate speech and noise streams, here we reduce the similarity between the conditions of the electrophysiological recording and the offline SIN behavioral task to a single overlapping feature: the presence of pitch-related information. Although several studies have not found a significant relationship between FFR-f0 measured in conditions of silence and SIN performance (e.g., Parbery-Clark et al., ), such relationships may be obscured by EEG-based FFR recordings which likely blend responses coming from different sources (Zhang and Gong, 2016; Tichko and Skoe, 2017). Despite the relative insensitivity of magnetoencephalography (MEG) to deep sources (which approximate radial sources, Baillet et al., ), sufficient information is preserved in the MEG signal for accurate localization of deeper structures such as the hippocampus, amygdala and thalamus (Attal and Schwartz, ; Dumas et al., ), and for the contributions from subcortical and cortical FFR generator sites to be separated (Coffey et al., ), which may increase the sensitivity of the experimental design to behavioral relationships. Therefore, an additional novel aspect of the present study is to use MEG to determine how FFR signals coming from distinct anatomical structures may be contributing to the putative relationship to SIN performance. We thus extend previous investigations of the anatomical origins of FFR signals to their behavioral meaning in the context of SIN perception.
If enhanced encoding is partly responsible for better auditory skills because a better quality signal is encoded from incoming sound and passed to higher-order cognitive processes and networks (Irvine, ; Musacchia et al., ), we would expect that the relationship between SIN and sound encoding would persist even under optimal listening conditions when the system is not challenged, and when the listener's attention is otherwise engaged. We therefore first measured SIN perception behaviorally, then in a separate session simultaneously recorded EEG and MEG data while listeners were presented with a speech sound in quiet as they watched a silent film. Secondarily, we evaluate correlations of measures of musical experience with FFR-f0 and SIN within our sample to evaluate the possible influence of musicianship on performance.
In addition to FFR-f0, which is derived from the higher-frequency EEG activity (Skoe and Kraus, 2010), other lower-frequency cortical potential measures covary with SIN performance, in particular the ERP P2 component (~200 ms post stimulus onset; Cunningham et al., ), which is also known to be related to speech processing and is sensitive to training effects (Key et al., ; Musacchia et al., ; Bidelman and Weiss, ; Tremblay et al., 2014). If the two signals represent sequential processes in the same processing stream, we would expect enhancements in FFR-f0 to be paralleled by enhancements in the strength of the ERP P2 component, and for each of these measures to be related to SIN accuracy. To test this hypothesis, we used distributed source modeling of the magnetic signals to localize the neural origins of the MEG FFR-f0 and the P2, and examined their spatial and statistical relationships to each other (as well as spatial relationships to preceding and following ERP components) for the first time in order to explore how these signals may be related. Collectively, these data should help us to understand the neural basis of inter-individual differences in sound encoding and its effects on the important real-world function of SIN perception.
Methods and materials
The experimental procedures concerning the MEG and (single channel, Cz) EEG recordings of the brain's response to the speech syllable /da/, and much of the pre-processing, have previously been reported in the context of determining their neural origins and will be discussed only briefly here (please see Coffey et al., “Methods” for details). The correlations between FFR-f0 strength and musicianship that are included in the summary of musical enhancements in Table 1 have been reported in Coffey et al.; all other findings have not been reported previously. Behavioral testing took place in a sound-attenuated room on different day prior to the MEG recording session.
Table 1
| Measure | Age of start | Practice hours |
|---|---|---|
| SIN | rs = −0.70, p = 0.006* | rs = 0.39, p = 0.10 |
| Fine pitch discrimination | rs = 0.45, p = 0.07 | rs = −0.67, p = 0.008* |
| FFR-f0 (right AC) | rs = −0.53, p = 0.05* | rs = 0.57, p = 0.04* |
| P2 amplitude | rs = −0.21, p = 0.25 | rs = 0.59, p = 0.05* |
Summary of evidence for musicianship-related behavioral and neurophysiological enhancements (N = 12).
Asterisks (*) indicate significant rank correlations (alpha < 0.05, one tailed). In general, earlier start ages and a larger number of practice hours are associated with enhancements, suggesting an influence of musical training. Note that lower fine pitch discrimination scores indicate better performance; therefore correlations in opposite directions are expected.
Participants
Data from the same 20 neurologically healthy young adults included in the previous study (Coffey et al., ) were included in this study (mean age: 25.7 years; SD = 4.2; 12 female; all were right-handed and had normal or corrected-to-normal vision; < = 25 dB hearing level thresholds for frequencies between 500 and 4,000 Hz assessed by pure-tone audiometry; and no history of neurological disorders). All but three subjects were native English speakers; the other three (one Korean, two French speakers) were highly proficient in English and all scored within the range of the native speakers on the HINT task, thus ruling out that any of our findings were due to second-language effects. Informed consent was obtained and all experimental procedures were approved by the Montreal Neurological Institute Research Ethics Board.
Speech-in-noise assessment
SIN was measured using a custom computerized implementation of the hearing in noise test (HINT; Nilsson, ) that allowed us to obtain a relative measure of SIN ability using a portable computer, without specialized equipment. In the standard HINT task, speech-spectrum noise is presented at a fixed level and sentences are varied in a staircase procedure to obtain a (single-value) SIN perceptual threshold (Nilsson, ). Our modified HINT task used a subset of the same sentence lists (Bench et al., ) and speech-spectrum noise, but presented thirty sentences in three empirically determined difficulty levels in randomized order: easy (2 dB SNR; i.e., target speech was 2 dB louder than noise), medium (−2 dB SNR), and difficult (−6 dB SNR). The sentences and noise were combined using sound processing software (Audacity, version 1.3.14-beta, http://audacity.sourceforge.net/; 44100Hz sampling frequency). Stimuli were presented diotically (i.e., identical speech and masker in each ear) via headphones (JVC HA-M5X) with the noise adjusted to a loud but not uncomfortable sound level on pilot subjects (~75 db SPL) and thereafter held constant. No verbal or visual feedback was given. A single overall accuracy score as the proportion of sentences correctly repeated back to experimenter was calculated by averaging the accuracy across all three levels; however, the score distributions showed a clear ceiling effect in the easiest level, with 12 out of 20 participants scoring over 95% (mean of subset of easy trials: 93.5%, SD = 7.5; medium trials: 79.9%, SD = 14.0; hard trials: 32.6%, SD = 12.1). We therefore excluded the easiest trials from the mean accuracy score in order to obtain a cleaner estimate of inter-individual variability; control analyses were conducted for the main research questions by assessing correlations between SIN scores at each level of difficulty and the FFR-f0 strength to ensure that the pattern of results was robust to this exclusion.
Fine pitch discrimination
Fine pitch discrimination thresholds were measured as described in Coffey et al. (), using a two-interval forced-choice task and a two-down one-up rule to estimate the threshold at 79% correct point on the psychometric curve (Levitt, ). The reference tone, which was presented once per trial, had a frequency of 500 Hz. The adaptive procedure was stopped after 15 reversals and the geometric mean of the last eight trials was recorded. Thresholds were derived from the average of five task repetitions.
Stimulus presentation
The stimulus for the MEG/EEG recordings was a 120-ms synthesized speech syllable (/da/) with a fundamental frequency in the sustained vowel portion of 98 Hz. The stimulus was presented binaurally at 80 dB SPL, ~14,000 times in alternating polarity, through Etymotic ER-3A insert earphones with foam tips (Etymotic Research). For five subjects, ~11,000 epochs were collected due to time constraints. Stimulus onset synchrony (SOA) was randomly selected between 195 and 205 ms from a normal distribution. A separate run was collected of ~600 stimulus repetitions spaced ~500 ms apart, to record later waves of the slower cortical responses. To control for attention and reduce fidgeting, a silent wildlife documentary (Yellowstone: Battle for Life, BBC, 2009) was projected onto a screen at a comfortable distance from the subject's face. This film was selected for being continuously visually appealing; subtitles were not provided in order to minimize saccades.
Neurophysiological recording and preprocessing
Two hundred and seventy-four channels of MEG (axial gradiometers), one channel of EEG data (Cz, 10–20 International System, averaged mastoid references), EOG and ECG, and one audio channel were simultaneously acquired using a CTF MEG System and its in-built EEG system (Omega 275, CTF Systems Inc.). All data were sampled at 12 kHz. Data preprocessing was performed with Brainstorm Tadel et al. (2011) and using custom Matlab scripts (The Mathworks Inc., MA, USA) as described in Coffey et al. (), and in brief, below.
FFR correlates of SIN accuracy
FFR-f0 strength was extracted from regions of interest (ROIs) in the auditory system (AC: auditory cortex, MGB: medial geniculate body of the thalamus, IC: inferior colliculus and CN: cochlear nucleus) using the MEG distributed source modeling approach described previously (for the specifications of each ROI, please see Methods in Coffey et al., ). Using this approach, the amplitude of a large set of dipoles are used to map activity originating in multiple generator sites; these are constrained by spatial priors derived from each subject's T1-weighted anatomical MRI scan (Baillet et al., ; Gross et al., ), from which cortical sources and subcortical structures were prepared using FreeSurfer (Fischl, ). As reported in Coffey et al., anatomical data were imported into Brainstorm (Tadel et al., 2011), and the brainstem and thalamic structures were combined with the cortex surface to form the image support of MEG distributed sources: the mixed surface/volume model included a triangulation of the cortical surface (~15,000 vertices), and brainstem and thalamus as a three-dimensional dipole grid (~18,000 points). An overlapping-sphere head model was computed for each run; this forward model explains how an electric current flowing in the brain would be recorded at the level of the sensors, with fair accuracy (Tadel et al., 2011). A noise covariance matrix was computed from 1-min empty-room recordings taken before each session. The inverse imaging model estimates the distribution of brain currents that account for data recorded at the sensors. We computed the MNE source distribution with unconstrained source orientations for each run using Brainstorm default parameters. The MNE source model is simple, robust to noise and model approximations, and very frequently used in literature (Hämäläinen, ). Source models for each run were averaged within subject. We extracted a timeseries of mean amplitude for each ROI and for each of the three orientations in the unconstrained orientation source model, from which three spectra were obtained by first windowing the signal (5 ms raised cosine ramp), zero padding to 1 s to enable a 1 Hz frequency resolution, with subsequent fast Fourier transform, and rescaling by the proportion of signal length to zero padding. The spectra of the three orientations were then summed in the frequency domain to obtain the amplitude of each subject's neurological response at the fundamental frequency, which was detected by an automatic script; this is referred hereafter as the FFR-f0 strength.
We first evaluated correlations between SIN accuracy scores and FFR-f0 strength averaged across bilateral pairs of structures, using Spearman's rho (rs; one-tailed). Non-parametric statistics were used throughout as FFR-f0 measures were generally not normally distributed (using Shapiro-Wilk's parametric hypothesis test of composite normality, the null hypothesis was rejected for AC, CN, and IC bilateral averages), and one-tailed tests were used as our goal was to test the specific hypothesis that higher amplitude FFR-would be related only to better behavioral performance, as stronger or less degraded FFRs in the presence of noise have been reported consistently in the EEG literature (Cunningham et al., ; Parbery-Clark et al., ,, ; Song et al., 2011), see also (Du et al., ). To correct for multiple comparisons, the false discovery rate (FDR) was controlled at δ = 0.05 where tests on multiple ROIs are used (Benjamini and Hochberg, ). The EEG equivalent of the FFR-f0 was also computed for comparison of sensitivity to behavioral measures. Correlations were computed between SIN accuracy and the left and right auditory cortex ROIs separately, as a lateralization effect in FFR-f0 strength and its relationship to measures of musicianship and fine pitch discrimination had been observed previously (see Figure 5c–e in Coffey et al., ). We tested for a stronger correlation on the right than left side using Fisher's r-to-Z transformation (one-tailed, alpha = 0.05).
Later cortical evoked responses
Event-related potentials (ERPs) within the 2–40 Hz band-pass filtered single-channel EEG data were obtained in order to establish a connection between previous FFR-ERP research that showed SIN sensitivity at ERP components P2 and N2 (Cunningham et al., ; Parbery-Clark et al., ) and the MEG data. In order to maximize the interpretability of the results with respect to a large body of work that has used EEG-based measures of P2 amplitude (which may not be entirely equivalent to their magnetic counterparts due to differences in each technique's sensitivity to source orientation), the primary measure of ERP amplitude is based on EEG rather than MEG; MEG sources were localized for the same time period.
Although, SIN perception has also been found to be correlated with N1 latency and amplitude measures (e.g., Parbery-Clark et al., ; Billings et al., ; Bidelman and Howell, ), N1 is strongly affected by the characteristics of the stimulus and its stimulation paradigm (Billings et al., ). In this paradigm, and using a single-EEG channel positioned at the vertex (Cz), we did not observe a clear N1 nor N2 from all subjects. We therefore took the amplitude of only P2 as a measure; this simpler metric occurs at a single time point and also allowed for a more straightforward comparison to and interpretation of the MEG equivalent. A researcher who was blinded to the subjects' FFR-f0 amplitudes and behavioral results at the time of measurement selected P2 wave peaks individually on ERP waves averaged across epochs for each subject (cortically processed; 2–40 Hz with −50 to 0 ms DC baseline correction; P2 was considered to be the strongest positive deflection within a ~40 ms window centered on the group grand average P2 at 183 ms). Amplitudes of these custom peaks were then correlated with SIN accuracy and FFR-f0 strength.
MEG evoked response fields (ERFs) on simultaneously recorded data were obtained in order to extend this work using distributed source modeling. The EEG cortical evoked response complex (ERP) elicited by the speech syllable /da/ consists of two positive waves at about 50–90 ms (“P1”) and between 170 and 200 ms (“P2” or “P1 prime”) and two negative waves at about 110 ms (“N1”) and after 200 ms (“N2” or “N1 prime”) (reviewed in Key et al., ); see also Cunningham et al., Figure 6 (Cunningham et al., ) and Musacchia et al., Figure 2 (Musacchia et al., ). For the purposes of this study we identified wave peaks in the ERP and ERF average at the group level for the SIN-sensitive P2 peak (183 ms), and at the earlier P1 component that has a well-known physiological origin in order to confirm the quality of data and validity of the analysis (60 ms; see Figures 3A,C).
Origins of later cortical ERP components
To confirm that the MEG data could be used to localize areas that showed above-baseline activity at the group level, and to observe the origins of the SIN-sensitive P2 wave in relation to preceding and following ERP waves, we first computed cortical volume MNE models based on each subject's T1-weighted MRI scan in which the orientation of sources was uncontrained, but their location was constrained within the volume encompassed by the cortical surface. These models were normalized to the baseline period (−50 to 0 ms). We exported 10 ms time windows around each peak of interest (mean-rectified signal amplitude) and for the baseline (−50 to 0 ms) for statistical analysis in the neuroimaging software package FSL (Smith et al., 2004; Jenkinson et al., ). These source volume maps were co-registered to the subject's high-resolution T1 anatomical MRI scan (FLIRT, 6 parameter linear transformation), and then to the 2 mm MNI152 template (12 parameter linear transformation, Evans et al., ). Normalized difference images were created by subtracting the baseline images from those of the peaks of interest and calculating z-scores within each image (P1 > Baseline, P2 > Baseline). Permutation testing was used to reveal locations where the magnetic signal was greater during peaks of interest as compared with baseline [non-parametric one-sample t-test (Winkler et al., 2014); 10,000 permutations]. The family-wise error rate was controlled using threshold-free cluster enhancement as implemented in FSL (p < 0.01), after applying a cortical mask of the MNI 152 template with the brainstem and cerebellum removed (these latter structures were not included in the MEG source model).
Comodulation of low and high frequency activity
We considered the spatial relationship between FFR-f0 generators and the source of the SIN-sensitive P2 wave by inspecting the FFR-f0 > Baseline and P2 > Baseline maps in the MEG data, and calculated Spearman's correlations between the FFR-f0 strength from each auditory cortex ROI (MEG) and the amplitude of the P2 wave measured with EEG.
Musicianship enhancements
Twelve subjects reported varying levels of musical experience on a range of musical instruments [primary instruments: piano (6), guitar (2), flute (1), saxophone (1), trumpet (1), voice (1)], as obtained by self-report using the Montreal Music History Questionnaire (Coffey et al., ). Start ages ranged from 5 to 12 years, and total cumulative practice hours ranged from 1,000 to 16,000 h. We assessed correlations between SIN accuracy and total music practice hours and age of training start. We then evaluated the relationship between P2 amplitude in the EEG recording and musicianship, and between fine pitch discrimination skills and musicianship.
Results
Behavioral scores
The mean averaged SIN score was 56.3% (SD = 12.0). Subjects with finer pitch discrimination ability had statistically better SIN accuracy (one-tailed rs = −0.47, p = 0.018).
MEG FFR-f0 strength is related to SIN throughout the auditory system
As described by Coffey et al. (), MEG is able to separate FFR activity arising from auditory cortex (Figure 1A), as well as brainstem and thalamus (Figure 1B). Relationships between FFR values measured via MEG from each ROI (averaged across left and right pairs) and SIN accuracy scores are presented in Figures 1C–F. A positive correlation between SIN accuracy and FFR-f0 strength was found at each of the four structures tested, statistically significant (FDR-corrected for multiple comparisons) in all but the inferior colliculus, where a similar trend was nonetheless noted. We did not find evidence of a relationship between the EEG-derived FFR-f0 and SIN accuracy (Figure 1G), nor did a relationship appear with the inclusion of age as a covariate (rs = 0.08, p = 0.37).
Figure 1
To confirm that the precautionary exclusion of the easiest SIN trials in which a ceiling effect was found was inconsequential with respect to the observed relationships between SIN and FFR-f0 strength, we recalculated the correlation between rAC FFR-f0 and SIN accuracy including all items (rs = 0.71, p = 0.0003; compare with the reported value with the exclusion, which is rs = 0.72, p < 0.0002). The general pattern of a positive correlation between rAC FFR-f0 and SIN accuracy was even replicated within the small subsets of easy (rs = 0.65, p = 0.001), medium (rs = 0.80, p < 0.0001), and hard items (rs = 0.57, p = 0.004), suggesting a robust relationship that is not highly sensitive to how the overall SIN accuracy score is calculated.
The relationship between SIN and cortical FFR-f0 is lateralized
The relationship between SIN and FFR-f0 strength from auditory cortical ROIs in each hemisphere is depicted in Figure 2. SIN accuracy was related to the strength of the FFR-f0 in both hemispheres, but was numerically larger on the right. We directly compared the strength of these correlations using Fisher's r-to-Z-transformation (one-tailed), and found it to be stronger in the right hemisphere (Z = −3.12, p = 0.001; the correlation between the FFR-f0 strength across two hemispheres, which is used for statistical comparison of correlation strength, was rs = 0.89).
Figure 2
Origins of later cortical ERP components
We confirmed that the MEG data analysis used here is suitable for localizing temporal lobe auditory activity at the group level using P1, which was known to originate in the primary auditory areas bilaterally (Figure 3D). Note that while we had selected the earliest maximum in the P1 wave in the EEG signal in order to capture primary auditory cortex activity (mean latency: 60 ms), the peak energy in the MEG signal is slightly later (~15 ms); nonetheless, visual inspection of the same analysis performed on a 10 ms window centered on 75 ms indicates that this analysis is not sensitive to minor variations in P1 window selection. The mean latency of P2, the second prominent positive EEG wave, was 183 ms (SD = 11 ms), and its mean amplitude was 4.1 uV (SD = 1.7). We confirmed, as previously reported by Cunningham et al. () using a pediatric sample, that P2 amplitude was related to SIN accuracy in the current sample (Figure 3B) thus providing a basis for further investigating FFR-f0 and P2 relationships. P1 amplitude was not related to SIN accuracy (rs = 0.11, p = 0.33). We then identified the sources of magnetic activity that was concurrent with the EEG-derived P2 wave, which proved to be relatively more anterior, and right-lateralized (Figure 3D; colored areas depict significant clusters corrected for multiple comparisons; maps are thresholded to best expose the areas of strongest signal).
Figure 3
Low and high frequency activity covary
Right but not left AC FFR-f0 strength was significantly related to P2 amplitude (Figures 4C,E). For completeness, we also calculated the correlation between the EEG FFR-f0 and P2 amplitude but it was not significant: rs = 0.23, p = 0.16). The magnetic equivalent of the P2 wave overlapped considerably with the FFR-f0 regions using a corrected significance-based threshold of p < 0.05. However, inspection of the centroid of each map showed that whereas the FFR-f0 sources were distributed in the posterior section of the superior temporal gyrus, the P2 wave's foci were more anterior (Figure 4D).
Figure 4
Measures of musicianship
Twelve out of 20 subjects reported some level of musical training. We previously showed that FFR-f0 strength in the right AC is related to hours of musical training and age of training start in this sample (Coffey et al.,
Discussion
In this study, we aimed to clarify whether individual differences observed in fundamental frequency (f0) encoding in the auditory system of normal-hearing adults is related to SIN perception. Toward this end, we filtered the neural responses in two frequency bands in order to isolate the higher-frequency 98 Hz f0 within the FFR, and the lower-frequency cortical responses (2–40 Hz), and we compared their strength, spatial origins, and relationships to behavioral and musical experience measures.
We first showed that the strength of the MEG-based FFR-f0 attributed to structures throughout the ascending auditory neuraxis, including the auditory cortex in each hemisphere, is positively correlated with SIN accuracy (Figures 1C–F), suggesting that basic periodic encoding is enhanced throughout the auditory system in people with better ability to perceive speech under challenging noise conditions. Although, the IC vs. SIN relationship falls short of significance, the trend is in the predicted direction (rs = 33, p = 0.075); it is therefore most parsimonious to assume that we failed to observe the relationship strongly here for reasons related to noise in the data rather than that the strength of the pitch representation is uncorrelated at a midpoint location between nuclei through which the same information passes.
Importantly, the principal similarity between the conditions under which the neurophysiological measurement was made (i.e., passive listening in silence) and the behavioral measure [i.e., deciphering speech in sentences, similar to the clinical HINT task (Nilsson,
According to such an approach, our neurophysiological recording paradigm did not explicitly engage auditory working memory and it did not require attention, which instead was directed to a silent film. Although, top-down processes might conceivably act spontaneously even in the presence of such a distraction, the speech sounds were presented in perfect clarity and did not offer more than one stream of information upon which top-down mechanisms of auditory stream analysis such as stream segregation (Başkent and Gaudrain,
Because we have eliminated cues that might be used by these higher-level processes that are known to affect SIN, either by enhancing the incoming signal (e.g., Lehmann and Schönwiesner,
Amplitude differences in the P2 cortical component and in FFR-f0 strength have been related to SIN perception differences between normal and learning-disordered children when recorded in noise (Cunningham et al.,
We found that variability in P2 amplitude correlated with inter-individual differences in SIN ability (Figure 3B) and in FFR-f0 strength (Figures 4C,E), despite that the responses were measured only in quiet conditions and in a normal healthy adult population. These results suggest that previously reported relationships may be present as a continuum in the population and even in optimal listening conditions. The MEG-FFR technique may allow us to more consistently observe behavioral and experience-related relationships with FFR-f0 strength in less challenging listening conditions as compared with the EEG-FFR. The EEG-FFR is likely a composite from several subcortical and cortical sources (Herdman et al.,
The spatial resolution of MEG source imaging may help to clarify the auditory processes that generate the ERP and ERF components, which have a long history in auditory neuroscience yet have predominantly been studied at the sensor level or using simpler models. We used distributed source modeling based on individual anatomy to localize the sources of each ERP/ERF wave (Figure 3D) with a view to confirming the localization of the wave of interest. The P1 wave originated bilaterally in the primary auditory areas, as expected (Liégeois-Chauvel et al.,
The relatively more anterior location of the P2 compared to the FFR generators (Figure 4D) could be explained by a right-lateralized anterior flow of pitch-relevant information that supports SIN processing, although future work will be needed to clarify whether the relationship between periodic encoding and later waves is causal in nature, or if these different frequency bands represent neural activity in parallel processing streams in neighboring neural populations. Furthermore, as mentioned previously, earlier ERP components such as N1 have previously been related to SIN perception (Billings et al.,
The strength of signal generators in the right hemisphere was stronger for both the FFR-f0 and P2 waves (Figures 4A,B,D). We found a positive correlation between P2 amplitude (measured with EEG at Cz) and MEG FFR-f0 in the right hemisphere (but not left; Figures 4C,E). These results corroborate previous work suggesting that the right auditory cortex is relatively specialized for pitch and tonal processing (Zatorre et al., 1994; Patel and Balaban,
Periodicity encoding is related to pitch information (Gockel et al.,
We propose that whether a musician advantage in SIN perception is observed or not in a given study may depend on interactions between the current state of the auditory system and the specific cognitive demands of the SIN task used in the study. Specifically, performance can depend on (1) the cues offered to the listener in the SIN paradigm [e.g., spatial cues and degree of information masking; (Swaminathan et al., 2015)]; (2) the degree to which an individual's experience has enhanced representations and mechanisms related to the available cues and caused them to be more strongly weighted; and (3) how well individuals can adapt to use alternative cues and mechanisms when one or more cues becomes less useful either through task differences like levels of noise (e.g., Du et al.,
Conclusion
In this study we present novel evidence that the quality of basic feed-forward periodicity encoding is related to the clinically relevant problem of separating speech from noise signals, and musical training. Specifically, in the absence of contextual cues and task demands and given a measurement tool that is sensitive to signal sources (i.e., MEG), enhancements in periodic sound encoding throughout the auditory neuraxis were correlated with better SIN ability in an offline task of sentence perception. This effect was observed to be stronger in the FFR signal localized to the right auditory cortex, and was related to slower cortical P2 wave amplitude measured by EEG, which is concurrent with activity in the right secondary auditory cortex measured with MEG and suggests an anterior flow of pitch-related information. Musicians show an advantage related to FFR strength, suggesting a possible role of experience. Our results suggest that inter-individual differences in neural correlates of basic periodic sound representation observed within the normal-hearing population (Ruggles et al.,
Statements
Ethics statement
This study was carried out in accordance with the recommendations of the Montreal Neurological Institute Research Ethics Board with written informed consent from all subjects. All subjects gave written informed consent in accordance with the Declaration of Helsinki. The protocol was approved by the Montreal Neurological Institute Research Ethics Board.
Author contributions
EC, SH, SB, and RZ designed the experiment, AC and EC collected the data, EC analyzed the data, and EC, SH, SB, and RZ wrote the paper.
Acknowledgments
We wish to thank Elizabeth Bock for assistance designing and testing the EEG–MEG recording set up, Marc Bouffard for help preparing subjects, Francois Tadel for his expert assistance with Brainstorm software, and Alexandre Lehmann for consultation regarding the interpretation of evoked auditory responses. The research was supported by operating grants to RZ from the Canadian Institutes of Health Research and from the Canada Fund for Innovation, by a Vanier Canada Graduate Scholarship to EC and by seed funding from the Centre for Research on Brain, Language and Music (CRBLM). SB was supported by the Killam Foundation, a Senior-Researcher grant from the Fonds de Recherche du Québec-Santé, a Discovery Grant from the Natural Science and Engineering Research Council of Canada and the National Institutes of Health (2R01EB009048-05).
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
References
1
AlainC.RoyeA.ArnottS. R. (2013). Middle- and long-latency auditory evoked potentials: what are they telling us on central auditory disorders, in Handbook of Clinical Neurophysiology: Disorders of Peripheral and Central Auditory Processing, Vol. 10, ed. CelesiaG. G. (Amsterdam: Elsevier B.V), 177–199.
2
AlainC.ZendelB. R.HutkaS.BidelmanG. M. (2014). Turning down the noise: the benefit of musical training on the aging auditory brain. Hear. Res.308, 162–173. 10.1016/j.heares.2013.06.008
3
AlbouyP.MattoutJ.BouetR.MabyE.SanchezG.AgueraP.-E.et al. (2013). Impaired pitch perception and memory in congenital amusia: the deficit starts in the auditory cortex. Brain136, 1639–1661. 10.1093/brain/awt082
4
AndersonS.KrausN. (2010a). Objective neural indices of speech-in-noise perception. Trends Amplif.14, 73–83. 10.1177/1084713810380227
5
AndersonS.KrausN. (2010b). Sensory-cognitive interaction in the neural encoding of speech in noise: a review. J. Am. Acad. Audiol.21, 575–585. 10.3766/jaaa.21.9.3
6
AndersonS.Parbery-ClarkA.White-SchwochT.KrausN. (2012). Aging affects neural precision of speech encoding. J. Neurosci.32, 14156–14164. 10.1523/JNEUROSCI.2176-12.2012
7
AndersonS.White-SchwochT.ChoiH. J.KrausN. (2013a). Training changes processing of speech cues in older adults with hearing loss. Front. Syst. Neurosci.7:97. 10.3389/fnsys.2013.00097
8
AndersonS.White-SchwochT.Parbery-ClarkA.KrausN. (2013b). A dynamic auditory-cognitive system supports speech-in-noise perception in older adults. Hear. Res.300, 18–32. 10.1016/j.heares.2013.03.006
9
AndohJ.MatsushitaR.ZatorreR. J. (2015). Asymmetric interhemispheric transfer in the auditory network: evidence from TMS, resting-state fMRI, and diffusion imaging. J. Neurosci.35, 14602–14611. 10.1523/JNEUROSCI.2333-15.2015
10
AttalY.SchwartzD. (2013). Assessment of subcortical source localization using deep brain activity imaging model with minimum norm operators: a MEG study. PLoS ONE8:e59856. 10.1371/journal.pone.0059856
11
BailletS.MosherJ. C.LeahyR. M. (2001). Electromagnetic brain mapping. IEEE Signal Process. Mag.18, 14–30. 10.1109/79.962275
12
BaşkentD.GaudrainE. (2016). Musician advantage for speech-on-speech perception. J. Acoust. Soc. Am.139, EL51–EL56. 10.1121/1.4942628
13
BenchJ.KowalA.BamfordJ. (1979). The BKB. (Bamford-Kowal-Bench) sentence lists for partially-hearing children. Br. J. Audiol.13, 108–112. 10.3109/03005367909078884
14
BendixenA. (2014). Predictability effects in auditory scene analysis: a review. Front. Neurosci.8:60. 10.3389/fnins.2014.00060
15
BenjaminiY.HochbergY. (1995). Controlling the false discovery rate: a practical and powerful approach to multiple controlling the false discovery rate: a practical and powerful approach to multiple testing. J. R. Stat. Soc.57, 289–300.
16
BidelmanG. M.KrishnanA.GandourJ. T. (2011a). Enhanced brainstem encoding predicts musicians' perceptual advantages with pitch. Eur. J. Neurosci.33, 530–538. 10.1111/j.1460-9568.2010.07527.x
17
BidelmanG. M.AlainC. (2015). Musical training orchestrates coordinated neuroplasticity in auditory brainstem and cortex to counteract age-related declines in categorical vowel perception. J. Neurosci.35, 1240–1249. 10.1523/JNEUROSCI.3292-14.2015
18
BidelmanG. M.HowellM. (2016). Functional changes in inter- and intra-hemispheric cortical processing underlying degraded speech perception. Neuroimage124, 581–590. 10.1016/j.neuroimage.2015.09.020
19
BidelmanG. M.GandourJ. T.KrishnanA. (2011b). Musicians and tone-language speakers share enhanced brainstem encoding but not perceptual benefits for musical pitch. Brain Cogn.77, 1–10. 10.1016/j.bandc.2011.07.006
20
BidelmanG. M.VillafuerteJ. W.MorenoS.AlainC. (2014). Age-related changes in the subcortical-cortical encoding and categorical perception of speech. Neurobiol. Aging35, 2526–2540. 10.1016/j.neurobiolaging.2014.05.006
21
BidelmanG.WeissM. (2014). Coordinated plasticity in brainstem and auditory cortex contributes to enhanced categorical speech perception in musicians. Eur. J. Neurosci.40, 2662–2672. 10.1111/ejn.12627
22
BillingsC. J.BennettK. O.MolisM. R.LeekM. R. (2011). Cortical encoding of signals in noise: effects of stimulus type and recording paradigm. Ear Hear.32, 53–60. 10.1097/AUD.0b013e3181ec5c46
23
BillingsC. J.McMillanG. P.PenmanT. M.GilleS. M. (2013). Predicting perception in noise using cortical auditory evoked potentials. J. Assoc. Res. Otolaryngol.14, 891–903. 10.1007/s10162-013-0415-y
24
BoebingerD.EvansS.RosenS.LimaC. F.ManlyT.ScottS. K. (2015). Musicians and non-musicians are equally adept at perceiving masked speech. J. Acoust. Soc. Am.137, 378–387. 10.1121/1.4904537
25
BregmanA. S. (1994). Auditory Scene Analysis: The Perceptual Organization of Sound. Boston, MA: MIT Press. Available online at: https://books.google.com/books?hl=en&lr=&id=jI8muSpAC5AC&pgis=1
26
BrokxJ. P. L.NooteboomS. G. (1982). Intonation and the perceptual separation of simultaneous voices. J. Phon.10, 23–36.
27
CarcagnoS.PlackC. J. (2011). Subcortical plasticity following perceptual learning in a pitch discrimination task. J. Assoc. Res. Otolaryngol.12, 89–100. 10.1007/s10162-010-0236-1
28
ChaK.ZatorreR. J.SchönwiesnerM. (2016). Frequency selectivity of voxel-by-voxel functional connectivity in human auditory cortex. Cereb. Cortex26, 211–224. 10.1093/cercor/bhu193
29
ChalikiaM. H.BregmanA. S. (1989). The perceptual segregation of simultaneous auditory signals: pulse train segregation and vowel segregation. Percept. Psychophys.46, 487–496. 10.3758/BF03210865
30
ChandrasekaranB.KrausN. (2010). The scalp-recorded brainstem response to speech: neural origins and plasticity. Psychophysiology47, 236–246. 10.1111/j.1469-8986.2009.00928.x
31
CoffeyE. B. J.HerholzS. C. (2013). Task decomposition: a framework for comparing diverse training models in human brain plasticity studies. Front. Hum. Neurosci.7:640. 10.3389/fnhum.2013.00640
32
CoffeyE. B. J.ColagrossoE. M. G.LehmannA.SchönwiesnerM.ZatorreR. J. (2016a). Individual differences in the frequency-following response: relation to pitch perception. Ed. Frederic Dick. PLoS ONE11:e0152374. 10.1371/journal.pone.0152374
33
CoffeyE. B. J.HerholzS. C.ChepesiukA. M. P.BailletS.ZatorreR. J. (2016b). Cortical contributions to the auditory frequency-following response revealed by MEG. Nat. Commun.7:11070. 10.1038/ncomms11070
34
CoffeyE. B. J.HerholzS. C.ScalaS.ZatorreR. J. (2011). Montreal Music History Questionnaire: a tool for the assessment of music-related experience in music cognition research, in The Neurosciences and Music IV: Learning and Memory, Conference (Edinburgh).
35
CoffeyE.MogileverN.ZatorreR. (2017). Speech-in-noise perception in musicians: a review. Hear. Res. 352, 49–69. 10.1016/j.heares.2017.02.006
36
CoffeyE. B. J.MusacchiaG.ZatorreR. J. (2016c). Cortical correlates of the auditory frequency-following and onset responses: EEG and fMRI evidence. J. Neurosci.37, 830–838. 10.1523/JNEUROSCI.1265-16.2016
37
CunninghamJ.NicolT.ZeckerS. G.BradlowA.KrausN. (2001). Neurobiologic responses to speech in noise in children with learning problems: deficits and strategies for improvement. Clin. Neurophysiol.112, 758–767. 10.1016/S1388-2457(01)00465-5
38
DuY.BuchsbaumB. R.GradyC. L.AlainC. (2014). Noise differentially impacts phoneme representations in the auditory and speech motor systems. Proc. Natl. Acad. Sci. U.S.A.111, 7126–7131. 10.1073/pnas.1318738111
39
DuY.BuchsbaumB. R.GradyC. L.AlainC. (2016). Increased activity in frontal motor cortex compensates impaired speech perception in older adults. Nat. Commun.7:12241. 10.1038/ncomms12241
40
DuY.KongL.WangQ.WuX.LiL. (2011). Auditory frequency-following response: a neurophysiological measure for studying the “cocktail-party problem”. Neurosci. Biobehav. Rev.35, 2046–2057. 10.1016/j.neubiorev.2011.05.008
41
DumasT.DubalS.AttalY.ChupinM.JouventR.MorelS.et al. (2013). MEG evidence for dynamic amygdala modulations by gaze and facial emotions. Ed. Andreas Keil. PLoS ONE8:e74145. 10.1371/journal.pone.0074145
42
EvansA. C.JankeA. L.CollinsD. L.BailletS. (2012). Brain templates and atlases. Neuroimage62, 911–922. 10.1016/j.neuroimage.2012.01.024
43
EvansS.MeekingsS.NuttallH. E.JasminK. M.BoebingerD.AdankP.et al. (2014). Does musical enrichment enhance the neural coding of syllables? Neuroscientific interventions and the importance of behavioral data. Front. Hum. Neurosci.8:964. 10.3389/fnhum.2014.00964
44
FischlB. (2012). FreeSurfer. Neuroimage62, 774–781. 10.1016/j.neuroimage.2012.01.021
45
FullerC. D.GalvinJ. J.MaatB.FreeR. H.BaşkentD. (2014). The musician effect: does it persist under degraded pitch conditions of cochlear implant simulations?Front. Neurosci.8:179. 10.3389/fnins.2014.00179
46
GockelH.CarlyonR.MehtaA.PlackC.LackC. H. J. P. (2011). The frequency following response. (FFR) may reflect pitch-bearing information but is not a direct representation of pitch. J. Assoc. Res. Otolaryngol.782, 767–782. 10.1007/s10162-011-0284-1
47
GolestaniN.RosenS.ScottS. K. (2009). Native-language benefit for understanding speech-in-noise: the contribution of semantics. Biling. Cogn.12, 385–392. 10.1017/S1366728909990150
48
GrossJ.BailletS.BarnesG. R.HensonR. N.HillebrandA.JensenO.et al. (2013). Good practice for conducting and reporting MEG research. Neuroimage65, 349–363. 10.1016/j.neuroimage.2012.10.001
49
HämäläinenM. S. (2009). MNE Software User's Guide v2.7. MGH/HMS/MIT Athinoula, A. Martinos Center for Biomedical Imaging Massachusetts General Hospital, Charlestown, MA.
50
HerdmanA. T.LinsO.RoonP.Van StapellsD. R.SchergM.PictonT. W.. (2002). Intracerebral sources of human auditory steady-state responses. Brain Topogr.15, 69–86. 10.1023/A:1021470822922
51
HerholzS. C.CoffeyE. B. J.PantevC.ZatorreR. J. (2015). Dissociation of neural networks for predisposition and for training-related plasticity in auditory-motor learning. Cereb. Cortex26, 3125–3134. 10.1093/cercor/bhv138
52
HydeK. L.PeretzI.ZatorreR. J. (2008). Evidence for the role of the right auditory cortex in fine pitch resolution. Neuropsychologia46, 632–639. 10.1016/j.neuropsychologia.2007.09.004
53
IrvineD. R. F. (1986). The Auditory Brainstem: a review of the structure and function of the auditory brainstem processing mechanisms, in Progress in Sensory Physiology, Vol. 7, eds AutrumH.OttosonD.PerlE. R.SchmidtR. F.ShimazuH.WillisW. D.OttosonD. (Berlin: Springer), 22–39.
54
JenkinsonM.BeckmannC. F.BehrensT. E. J.WoolrichM. W.SmithS. M. (2012). FSL. Neuroimage62, 782–790. 10.1016/j.neuroimage.2011.09.015
55
KeyA. P. F.DoveG. O.MaguireM. J. (2005). Linking brainwaves to the brain: an ERP primer. Dev. Neuropsychol.27, 183–215. 10.1207/s15326942dn2702_1
56
KingA.HopkinsK.PlackC. J. (2016). Differential group delay of the frequency following response measured vertically and horizontally. J. Assoc. Res. Otolaryngol.17, 133–143. 10.1007/s10162-016-0556-x
57
KrausN.SlaterJ.ThompsonE. C.HornickelJ.StraitD. L.NicolT.et al. (2014). Music enrichment programs improve the neural encoding of speech in at-risk children. J. Neurosci.34, 11913–11918. 10.1523/JNEUROSCI.1881-14.2014
58
KrausN.StraitD. L.Parbery-ClarkA. (2012). Cognitive factors shape brain networks for auditory skills: spotlight on auditory working memory. Ann. N. Y. Acad. Sci.1252, 100–107. 10.1111/j.1749-6632.2012.06463.x
59
KurikiS.KandaS.HirataY. (2006). Effects of musical experience on different components of MEG responses elicited by sequential piano-tones and chords. J. Neurosci.26, 4046–4053. 10.1523/JNEUROSCI.3907-05.2006
60
KuwadaS.AndersonJ. (2002). Sources of the scalp-recorded amplitude-modulation following response. J. Am. Acad. Audiol.13, 188–204.
61
LaineM.RinneJ. O.KrauseB. J.TeräsM.SipiläH. (1999). Left hemisphere activation during processing of morphologically complex word forms in adults. Neurosci. Lett.271, 85–88. 10.1016/S0304-3940(99)00527-3
62
LappeC.TrainorL. J.HerholzS. C.PantevC. (2011). Cortical plasticity induced by short-term multimodal musical rhythm training. PLoS ONE6:e21493. 10.1371/journal.pone.0021493
63
LehmannA.SchönwiesnerM. (2014). Selective attention modulates human auditory brainstem responses: relative contributions of frequency and spatial cues. PLoS ONE9:e85442. 10.1371/journal.pone.0085442
64
LevittH. (1971). Transformed up-down methods in psychoacoustics. J. Acoust. Soc. Am.49, 467–477. 10.1121/1.1912375
65
Liégeois-ChauvelC.MusolinoA.BadierJ. M.MarquisP.ChauvelP. (1994). Evoked potentials recorded from the auditory cortex in man: evaluation and topography of the middle latency components. Electroencephalogr. Clin. Neurophysiol.92, 204–214. 10.1016/0168-5597(94)90064-7
66
LombardE. (1911). Le signe de l'élevation de la voix. Ann Mal Oreille Larynx Nez Pharynx37, 101–119.
67
MathysC.LouiP.ZhengX.SchlaugG. (2010). Non-invasive brain stimulation applied to Heschl's gyrus modulates pitch discrimination. Front. Psychol.1:193. 10.3389/fpsyg.2010.00193
68
MooreB. C. J.GockelH. (2002). Factors influencing sequential stream segregation. Acta Acust. United Acust.88, 320–332. Available online at: http://www.ingentaconnect.com/content/dav/aaua/2002/00000088/00000003/art00004
69
MusacchiaG.SamsM.SkoeE.KrausN. (2007). Musicians have enhanced subcortical auditory and audiovisual processing of speech and music. Proc. Natl. Acad. Sci. U.S.A.104, 15894–15898. 10.1073/pnas.0701498104
70
MusacchiaG.StraitD. L.KrausN. (2008). Relationships between behavior, brainstem and cortical encoding of seen and heard speech in musicians and non-musicians. Hear. Res.241, 34–42. 10.1016/j.heares.2008.04.013
71
NäätänenR.PictonT. (1987). The N1 wave of the human electric and magnetic response to sound: a review and an analysis of the component structure. Psychophysiology24, 375–425. 10.1111/j.1469-8986.1987.tb00311.x
72
NilssonM. (1994). Development of the Hearing In Noise Test for the measurement of speech reception thresholds in quiet and in noise. J. Acoust. Soc. Am.95, 1085–1099. 10.1121/1.408469
73
NygaardL. C.SommersM. S.PisoniD. B. (1994). Speech perception as a talker-contingent process. Psychol. Sci.5, 42–46. 10.1111/j.1467-9280.1994.tb00612.x
74
Parbery-ClarkA.AndersonS.HittnerE.KrausN. (2012a). Musical experience strengthens the neural representation of sounds important for communication in middle-aged adults. Front. Aging Neurosci.4:30. 10.3389/fnagi.2012.00030
75
Parbery-ClarkA.AndersonS.HittnerE.KrausN. (2012b). Musical experience offsets age-related delays in neural timing. Neurobiol Aging33, 1483.e1–1483.e4. 10.1016/j.neurobiolaging.2011.12.015
76
Parbery-ClarkA.MarmelF.BairJ.KrausN. (2011a). What subcortical-cortical relationships tell us about processing speech in noise. Eur. J. Neurosci.33, 549–557. 10.1111/j.1460-9568.2010.07546.x
77
Parbery-ClarkA.SkoeE.KrausN. (2009a). Musical experience limits the degradative effects of background noise on the neural processing of sound. J. Neurosci.29, 14100–14107. 10.1523/JNEUROSCI.3256-09.2009
78
Parbery-ClarkA.SkoeE.LamC.KrausN. (2009b). Musician enhancement for speech-in-noise. Ear Hear.30, 653–661. 10.1097/AUD.0b013e3181b412e9
79
Parbery-ClarkA.StraitD. L.KrausN. (2011c). Context-dependent encoding in the auditory brainstem subserves enhanced speech-in-noise perception in musicians. Neuropsychologia49, 3338–3345. 10.1016/j.neuropsychologia.2011.08.007
80
Parbery-ClarkA.StraitD. L.AndersonS.HittnerE.KrausN. (2011b). Musical experience and the aging auditory system: implications for cognitive abilities and hearing speech in noise. PLoS ONE6:e18082. 10.1371/journal.pone.0018082
81
PatelA. D. (2014). Can nonlinguistic musical training change the way the brain processes speech? The expanded OPERA hypothesis. Hear. Res.308, 98–108. 10.1016/j.heares.2013.08.011
82
PatelA. D.BalabanE. (2001). Human pitch perception is reflected in the timing of stimulus-related cortical activity. Nat. Neurosci.4, 839–844. 10.1038/90557
83
PattersonR. D.UppenkampS.JohnsrudeI. S.GriffithsT. D. (2002). The processing of temporal pitch and melody information in auditory cortex. Neuron36, 767–776. 10.1016/S0896-6273(02)01060-7
84
PickeringM. J.GarrodS. (2007). Do people use language production to make predictions during comprehension?Trends Cogn. Sci.11, 105–110. 10.1016/j.tics.2006.12.002
85
PressnitzerD.SuiedC.ShammaS. (2011). Auditory scene analysis: the sweet music of ambiguity. Front. Hum. Neurosci.5:158. 10.3389/fnhum.2011.00158
86
RinneT.BalkM. H.KoistinenS.AuttiT.AlhoK.SamsM. (2008). Auditory selective attention modulates activation of human inferior colliculus. J. Neurophysiol.100, 3323–3327. 10.1152/jn.90607.2008
87
RossB.FujiokaT. (2016). 40-Hz oscillations underlying perceptual binding in young and older adults. Psychophysiology53, 974–990. 10.1111/psyp.12654
88
RossB.MiyazakiT.FujiokaT. (2012). Interference in dichotic listening: the effect of contralateral noise on oscillatory brain networks. Eur. J. Neurosci.35, 106–118. 10.1111/j.1460-9568.2011.07935.x
89
RugglesD. R.FreymanR. L.OxenhamA. J. (2014). Influence of musical training on understanding voiced and whispered speech in noise. PLoS ONE9:e86980. 10.1371/journal.pone.0086980
90
RugglesD.BharadwajH.Shinn-CunninghamB. G. (2011). Normal hearing is not enough to guarantee robust encoding of suprathreshold features important in everyday communication. Proc. Natl. Acad. Sci. U.S.A.108, 15516–15521. 10.1073/pnas.1108912108
91
SchneiderP.SchergM.DoschH. G.SpechtH. J.GutschalkA.RuppA. (2002). Morphology of Heschl's gyrus reflects enhanced activation in the auditory cortex of musicians. Nat. Neurosci.5, 688–694. 10.1038/nn871
92
ShahinA.BosnyakD. J.TrainorL. J.RobertsL. E. (2003). Enhancement of neuroplastic P2 and N1c auditory evoked potentials in musicians. J. Neurosci.23, 5545–5552. Available online at: http://www.jneurosci.org/content/23/13/5545
93
ShtyrovY.KujalaT.AhveninenJ.TervaniemiM.AlkuP.IlmoniemiR. J.et al. (1998). Background acoustic noise and the hemispheric lateralization of speech processing in the human brain: magnetic mismatch negativity study. Neurosci. Lett.251, 141–144. 10.1016/S0304-3940(98)00529-1
94
ShtyrovY.KujalaT.IlmoniemiR. J.NaR.NenÈ. (1999). Noise Affects Speech- Signal Processing Differently in the Cerebral Hemispheres. Available online at: https://insights.ovid.com/neuroreport/nerep/1999/07/130/noise-affects-speech-signal-processing-differently/34/00001756.
95
SkoeE.KrausN. (2010). Auditory brain stem response to complex sounds: a tutorial. Ear Hear.31, 302–324. 10.1097/AUD.0b013e3181cdb272
96
SlaterJ.SkoeE.StraitD. L.O'ConnellS.ThompsonE.KrausN. (2015). Music training improves speech-in-noise perception: Longitudinal evidence from a community-based music program. Behav. Brain Res.291, 244–252. 10.1016/j.bbr.2015.05.026
97
SmithS. M.JenkinsonM.WoolrichM. W.BeckmannC. F.BehrensT. E. J.Johansen-BergH.et al. (2004). Advances in functional and structural MR image analysis and implementation as FSL. Neuroimage23(Suppl. 1), S208–S219. 10.1016/j.neuroimage.2004.07.051
98
SongJ. H.SkoeE.BanaiK.KrausN. (2011). Perception of speech in noise: neural correlates. J. Cogn. Neurosci.23, 2268–2279. 10.1162/jocn.2010.21556
99
SongJ. H.SkoeE.BanaiK.KrausN. (2012). Training to improve hearing speech in noise: biological mechanisms. Cereb. Cortex22, 1180–1190. 10.1093/cercor/bhr196
100
SongJ.SkoeE.WongP.KrausN. (2008). Plasticity in the adult human auditory brainstem following short-term linguistic training. J. Cogn. Neurosci.20, 1892–1902. 10.1162/jocn.2008.20131
101
SouzaP.GehaniN.WrightR.McCloyD. (2013). The advantage of knowing the talker. J. Am. Acad. Audiol.24, 689–700. 10.3766/jaaa.24.8.6
102
StraitD. L.KrausN. (2011). Can you hear me now? Musical training shapes functional brain networks for selective auditory attention and hearing speech in noise. Front. Psychol.2:113. 10.3389/fpsyg.2011.00113
103
StraitD. L.KrausN.Parbery-ClarkA.AshleyR. (2010). Musical experience shapes top-down auditory mechanisms: evidence from masking and auditory attention performance. Hear. Res.261, 22–29. 10.1016/j.heares.2009.12.021
104
StraitD. L.Parbery-ClarkA.HittnerE.KrausN. (2012). Musical training during early childhood enhances the neural encoding of speech in noise. Brain Lang.123, 191–201. 10.1016/j.bandl.2012.09.001
105
SugaN. (2012). Tuning shifts of the auditory system by corticocortical and corticofugal projections and conditioning. Neurosci. Biobehav. Rev.36, 969–988. 10.1016/j.neubiorev.2011.11.006
106
SuiedC.BonneelN.Viaud-DelmonI. (2009). Integration of auditory and visual information in the recognition of realistic objects. Exp. Brain Res.194, 91–102. 10.1007/s00221-008-1672-6
107
SummerfieldQ.AssmannP. F. (1999). Perception of concurrent vowels: effects of harmonic misalignment and pitch-period asynchrony. J. Acoust. Soc. Am.89, 1364–1377.
108
SwaminathanJ.MasonC. R.StreeterT. M.BestV.KiddG.PatelA. D. (2015). Musical training, individual differences and the cocktail party problem. Sci. Rep.5, 1–10. 10.1038/srep11628
109
TadelF.BailletS.MosherJ. C.PantazisD.LeahyR. M. (2011). Brainstorm: a user-friendly application for MEG/EEG analysis. Comput. Intell. Neurosci.2011, 1–13. 10.1155/2011/879716
110
TichkoP.SkoeE. (2017). Frequency-dependent fine structure in the frequency-following response: the byproduct of multiple generators. Hear. Res.348, 1–15. 10.1016/j.heares.2017.01.014
111
TierneyA.KrizmanJ.SkoeE.JohnstonK.KrausN. (2013). High school music classes enhance the neural processing of speech. Front. Psychol.4:855. 10.3389/fpsyg.2013.00855
112
TremblayK. L.RossB.InoueK.McClannahanK.ColletG. (2014). Is the auditory evoked P2 response a biomarker of learning?Front. Syst. Neurosci.8:28. 10.3389/fnsys.2014.00028
113
VargheseL.BharadwajH. M.Shinn-CunninghamB. G. (2015). Evidence against attentional state modulating scalp-recorded auditory brainstem steady-state responses. Brain Res.1626, 146–164. 10.1016/j.brainres.2015.06.038
114
WeissM. W.BidelmanG. M. (2015). Listening to the brainstem: musicianship enhances intelligibility of subcortical representations for speech. J. Neurosci.35, 1687–1691. 10.1523/JNEUROSCI.3680-14.2015
115
WilsonR. H.McArdleR. A.SmithS. L. (2007). An evaluation of the BKB-SIN, HINT, QuickSIN, and WIN materials on listeners with normal hearing and listeners with hearing loss. J. speech Lang. Hear. Res.50, 844–856. 10.1044/1092-4388(2007/059)
116
WinklerA. M.RidgwayG. R.WebsterM. A.SmithS. M.NicholsT. E. (2014). Permutation inference for the general linear model. Neuroimage92, 381–397. 10.1016/j.neuroimage.2014.01.060
117
ZatorreR. J.BelinP.PenhuneV. B. (2002). Structure and function of auditory cortex: music and speech. Trends Cogn. Sci.6, 37–46. 10.1016/S1364-6613(00)01816-7
118
ZatorreR. J.EvansA.MeyerE. (1994). Neural mechanisms underlying melodic perception and memory for pitch. J. Neurosci.14, 1908–1919.
119
ZendelB. R.AlainC. (2012). Musicians experience less age-related decline in central auditory processing. Psychol. Aging27, 410–417. 10.1037/a0024816
120
ZhangX.GongQ. (2016). Correlation between the frequency difference limen and an index based on principal component analysis of the frequency-following response of normal hearing listeners. Hear. Res. 344, 255–264. 10.1016/j.heares.2016.12.004
121
Zion GolumbicE.CoganG. B.SchroederC. E.PoeppelD. (2013). Visual input enhances selective speech envelope tracking in auditory cortex at a “cocktail party”. J. Neurosci.33, 1417–1426. 10.1523/JNEUROSCI.3675-12.2013
Summary
Keywords
frequency-following response, speech-in-noise, magnetoencephalography, electroencephalography, neuroplasticity, auditory perception, inter-individual variability, musical training
Citation
Coffey EBJ, Chepesiuk AMP, Herholz SC, Baillet S and Zatorre RJ (2017) Neural Correlates of Early Sound Encoding and their Relationship to Speech-in-Noise Perception. Front. Neurosci. 11:479. doi: 10.3389/fnins.2017.00479
Received
25 May 2017
Accepted
11 August 2017
Published
25 August 2017
Volume
11 - 2017
Edited by
Claude Alain, Rotman Research Institute, Canada
Reviewed by
Gavin M. Bidelman, University of Memphis, United States; Yang Zhang, University of Minnesota, United States
Updates

Check for updates
Copyright
© 2017 Coffey, Chepesiuk, Herholz, Baillet and Zatorre.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Emily B. J. Coffey emily.coffey2@mail.mcgill.ca
This article was submitted to Auditory Cognitive Neuroscience, a section of the journal Frontiers in Neuroscience
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.