Open APA research · Launch collection CC BY 3.0

Domain-general mechanisms for speech segmentation: The role of duration information in language learning.

Frost RL, Monaghan P, Tatsumi T.

Journal of experimental psychology. Human perception and performanceAmerican Psychological Association2016-11-28DOI 10.1037/xhp0000325

Abstract

Speech segmentation is supported by multiple sources of information that may either inform language processing specifically, or serve learning more broadly. The Iambic/Trochaic Law (ITL), where increased duration indicates the end of a group and increased emphasis indicates the beginning of a group, has been proposed as a domain-general mechanism that also applies to language. However, language background has been suggested to modulate use of the ITL, meaning that these perceptual grouping preferences may instead be a consequence of language exposure. To distinguish between these accounts, we exposed native-English and native-Japanese listeners to sequences of speech (Experiment 1) and nonspeech stimuli (Experiment 2), and examined segmentation using a 2AFC task. Duration was manipulated over 3 conditions: sequences contained either an initial-item duration increase, or a final-item duration increase, or items of uniform duration. In Experiment 1, language background did not affect the use of duration as a cue for segmenting speech in a structured artificial language. In Experiment 2, the same results were found for grouping structured sequences of visual shapes. The results are consistent with proposals that duration information draws upon a domain-general mechanism that can apply to the special case of language acquisition. (PsycINFO Database Record

Attribution and reuse record

Authors
Frost RL, Monaghan P, Tatsumi T.
Original journal
Journal of experimental psychology. Human perception and performance
Publisher
American Psychological Association
Publication date
2016-11-28
DOI
10.1037/xhp0000325
License
CC BY 3.0
Open repository
Europe PMC · PMC5327892
Collection
School leadership launch collection

Presented by the Journal for School Superintendents under the license identified in the article’s open full-text record. The original authors and publisher do not endorse this journal or its agent.

Open full text

Read the scholarly record

Experiment 1: Use of Duration Information in Speech Segmentation

In this experiment, we replicated previous studies of native-English listeners’ use of a final-syllable duration increase cue for segmenting artificial speech, and extended this test to a group of native-Japanese listeners. If use of final syllable duration increase is language dependent then we expect a smaller effect for Japanese than English speakers. If the effect is domain general, and not learned from language exposure, then we expect similar use of the cue by both English and Japanese speakers.

From power analyses based on Saffran et al.’s (1996) study of duration cues, we determined that 12 participants per condition would result in power of .79 for finding a difference between any two conditions (effect size in Saffran et al.’s (1996) study resulted in Cohen’s d = 1.113 for the comparison between their final duration increase and no duration increase conditions). For the English group, 36 students from Lancaster University, 11 males and 25 females with a mean age of 20.69 years ( SD = 4.06), volunteered to participate in the study for course credit. All participants reported English as their first language and reported no hearing or vision problems. For the Japanese group, 34 native Japanese listeners who were students and staff at the Tokyo University of Foreign Studies, 11 males and 23 females with a mean age of 22.69 years ( SD = 4.655), volunteered to take part in the study, and received 700 yen for their participation. Data for a further two participants were collected, but were removed from the analysis because they were outliers in terms of age (ages 61 and 59). The Japanese listeners all had some knowledge of a second language, required as part of the high-school curriculum, however, we did not collect information on level of proficiency in other languages for the participants. Although experience with a second language is likely to have a very limited effect on processing here (see Boll-Avetisyan et al., 2015 , and Molnar et al., 2014 ), we return to the issue of influence of second language learning on performance in the General Discussion.

Materials

We constructed an artificial language from six consonants (/b, d, g, k, p, t/), each used twice, and four vowels (in English:/æ, i, ɔ, u/, in Japanese:/a, i, o, ɯ/), each used three times, which were combined to produce 12 distinct CV syllables. The consonants were selected as those that were attested in both English and Japanese speech, and plosives were selected because these had a distinctive onset enabling duration of the syllable or mora to be processed by the listener. Vowels were selected to ensure distinctiveness in the productions of the speech synthesizer by varying both height and position of the vowels. The syllables were then concatenated to create four trisyllabic words (e.g., bogada, dugibu, kitapo, pikotu ), which were compiled pseudorandomly into a speech stream, with the restriction that no word was directly repeated. Within words, transitional probabilities between syllables were 1, and between words transitional probabilities were .33. Four different versions of the language were generated and counterbalanced across participants, to ensure that no biases for particular sequences influenced participants’ performance ( Onnis, Monaghan, Richmond, & Chater, 2005 ). These versions had different combinations of consonants and vowels within syllables or mora, and different combinations of syllables/mora comprising the four trisyllabic words. We ensured that words in the experimental languages were not preexisting words in English or Japanese.

For the English listeners, speech was synthesized using the Festival speech synthesizer ( Black, Taylor, & Caley, 1990 ), using the kal British English diphone database. For the Japanese listeners, speech was synthesized with MBROLA ( Dutoit, 1997 ) using the jp1 diphone database. For the synthesized speech, the diphone database permitted allophonic variation to be present, to result in more naturalistic speech. Duration increases were implemented in the training speech only. Duration of syllables was 233 ms, and 333 ms for increased duration syllables, similar to the speech used by Saffran et al. (1996) . Duration increase was implemented by increasing the vowel duration during synthesis by 100 ms. Speech was produced in a monotone with mean F0 of 120 Hz. Continuous speech was produced from streams of 150 words, in one of three conditions: initial-syllable duration increase (ISI), where the first syllable in every word was increased in duration; no duration increase (NI), where every syllable had equal duration; and final-syllable duration increase (FSI), where the third syllable in each word was extended in duration. There were no pauses between any of the syllables in the speech. Training time was 120 s for the initial and final syllable duration increase conditions, and 105 s for the no duration increase condition, and speech streams were edited to fade in and out for the first and last 5 s.

Critically, for the test materials the four word stimuli were synthesized with equal syllable duration (all 233 ms), regardless of the training condition. This was so that we could determine how the duration information was used to detect the structure, similar to previous speech segmentation studies applying syllable duration variation ( Saffran et al., 1996 ). Including the duration variation during testing would have meant that responses could be driven entirely by preferences for particular sequences regardless of the structure of the language. Sixteen part-words were also generated with equal syllable durations of 233 ms. Part-words occurred in the training speech but straddled word boundaries, comprising the last syllable of one word and the first two syllables of another word (so, for words of the form ABC, part-words would be of the form CAB), or the last two syllables of one word and the first syllable of another (BCA). Table 1 shows the relationship between the duration cue’s position during training for the word and part-word test items.

Procedure

Participants within each language group were randomly assigned to one of three duration conditions. All instructions were presented in the listeners’ native language. Instructions given in Japanese were based on direct translations of instructions given to the English listeners (performed by a native speaker of Japanese) to ensure precise comparability. Participants were instructed as follows: “Listen to the speech and try to determine the structure of the language.” After exposure to the speech, participants completed a forced-choice test containing 16 trials. For each trial, participants listened to a word and a part-word, separated by a 1-s pause and were instructed to “Select which sequence best matches the language you have just heard,” giving a key-press response of “1” for the first or “2” for the second sequence. Order of words and part-words within pairs was counterbalanced. In the test, all words occurred 4 times, and each part-word occurred once. Training and testing was then repeated, to examine effects of repetition of words, and to assess whether any additional learning took place over the course of the test phase. If performance improved at the second test then this could indicate that exposure to words and part-words affected performance on the task. If there was no effect then this means that the test itself was not affecting participants’ processing of individual stimuli. The experiment lasted for approximately 8 min. All participants were tested individually in an isolated booth, and received training and test items through closed-cup headphones. Participants listened to the speech at a volume that they found comfortable.

Results

To test the overall effect of duration on performance, we conducted a repeated-measures ANOVA on accuracy scores (proportion of selections of words over part-words), with language group (English or Japanese) and cue condition (ISI, NI, or FSI) as between subjects factors, and test time (first and second test) as a within subjects factor.

There was no significant effect of test time (Test 1: M = .652, SE = .025; Test 2: M = .678, SE = .026, F (1, 64) = 2.184, p = .144, η p 2 = .033). Interactions involving test time were not significant and are not further reported (all p > .05). This indicated that no learning took place during the testing.

There was a significant effect of cue condition, F (2, 64) = 8.323, p = .001, η p 2 = .206, and a linear contrast respecting the predicted order of cues (ISI, NI, FSI) was highly significant, F (2, 64) = 8.323, p = .001, η p 2 = .206 (see Figure 1 ). Dunnett’s post hoc t tests conducted to respect the linear contrast order revealed that the FSI condition ( M = .781, SE = .040) resulted in more accurate word identification from the speech than the NI condition ( M = .638, SE = .030), p = .009, and the ISI condition ( M = .573, SE = .040), p < .001. The NI and ISI conditions did not differ significantly, p = .199.

Mean word-identification score for participants in each cue condition, given for English and Japanese listeners. Error bars show ±1 SEM . See the online article for the color version of this figure.

There was a significant effect of language group, F (1, 64) = 4.321, p = .042, η p 2 = .063, with Japanese listeners ( M = .712, SE = .039) more accurate than English listeners ( M = .621, SE = .026) at identifying words from speech. There was no significant interaction between language group and cue condition, F (2, 64) = .019, p = .981.

Discussion

For English listeners, the results replicated previous studies demonstrating the benefit of final-syllable duration increase for identifying word boundaries in speech, where word boundaries are defined in terms of transitional probabilities between syllables ( Saffran et al., 1996 ). The linear effect of cue condition showed that increasing the duration of the initial syllable had a slight detrimental effect on identifying word boundaries compared with the other conditions (though this was not significantly different than chance), whereas increasing the duration of the final syllable meant that such boundaries were isolated more accurately compared with when no duration cue was present in the speech.

Interestingly, the results for Japanese listeners showed a very similar pattern. As with the English listeners, increasing the duration of the final syllable improved accuracy for word identification, where word boundaries were defined by the same transitional probabilities present in the speech heard by native-English listeners. Together with the fact that there were no interactions concerning native language and cue condition, this finding indicates that both language groups were using duration information in the same way to segment speech, contrary to findings of previous comparisons of English and Japanese listeners for unstructured sequences ( Iversen et al., 2008 ; Kusumoto & Moreton, 1997 ; Yoshida et al., 2010 ).

A key difference between our study and previous comparisons of English and Japanese listeners is the statistical structure of the stimuli. For Iversen et al. (2008) , Kusumoto and Moreton (1997) , and Yoshida et al. (2010) , participants listened to pairs of tones or syllables, with no transitional probability information. In our study, transitional probabilities were varied, such that there was a statistical structure to the speech to be discovered by participants. This could have been sufficient to guide the sporadic use of either an iambic or a trochaic grouping preference seen for Iversen et al.’s (2008) Japanese listeners toward a preference for grouping based on duration increase in the final element (to reflect structure). For unstructured sequences, such as those in Iversen et al.’s (2008) study, it is just not possible to integrate the prosodic information with the statistical information present in the speech potentially obscuring use of a domain-general iambic preference for detecting structure.

Importantly, the presence of structure in the speech did not lead English or Japanese listeners to prioritize a grouping preference for final duration increase regardless of the informational structure of the speech. Participants hearing speech with increased duration of the initial syllable demonstrated a non-significant deficit in word identification compared with those in the no duration increase condition, but performance was not lower than chance, meaning that the duration cue did not entirely override the statistical information; rather, both statistical and prosodic cues were used in combination. Thus, duration-related perceptual grouping is used in conjunction with, and not as a consequence of, the statistical structure of the speech.

Another important methodological distinction between our study and previous cross-linguistic comparisons of use of duration information in speech was the fact that in our study no durational information was present in the stimuli during testing. Rather, we tested learning as a consequence of using duration information, instead of testing the immediate influence of the duration cue on sequence perception. This enables us to determine how duration can be utilized to support learning of sequential structure, and avoids effects of perceptual capture during testing.

Our results are consistent with the idea that the use of duration variation as a cue is independent of language exposure; both English and Japanese listeners used final syllable duration increase to a similar degree, regardless of the magnitude of this effect in their background language. The results are thus indicative of duration information being available as a domain general cue. The next experiment tested whether duration variation exerted a similar effect for grouping visual sequences of shapes for English and Japanese listeners, to determine whether the effect of duration is specific to language stimuli, or is generalizable across modalities. If the use of duration information is language specific (as would be expected if the iambic preference is learned from language structure), then we would not find an effect of duration on grouping of structured sequences of shapes. However, if the preference is modality independent (as would be expected from a domain-general mechanism that is not learned as a consequence of language exposure), then the results should mirror those of the language stimuli in Experiment 1.

Materials

The materials were identical for both language groups, and were constructed to match the structural properties of the language used in Experiment 1, with each syllable being replaced by a shape. We selected 12 geometric shapes printed in black on a gray background, taken from Fiser and Aslin (2002) , each 170 × 170 pixels in size. Figure 2 shows some example stimuli used in the study. Shapes were concatenated into four triplets, with four different random arrangements counterbalanced across participants to ensure that no biases in terms of sequence preferences adversely affected the results. For the shapes, duration increase was implemented by displaying shapes on screen for 100 ms longer than the other shapes. Shapes with standard duration appeared at the center of a computer screen for 225 ms, and shapes with increased duration appeared for 325 ms. We wanted to ensure that the durational increases implemented for shapes were similar to those implemented for the speech stimuli, and while Peña et al. (2011) varied length of shape stimuli from 320 ms to 800 ms, such large duration differences could have resulted in divergence of processing from the speech stimuli. Though the overall duration is controlled between the speech and shape stimuli, it is important to note that increased duration of the speech resulted in a change across the vowel, which is not a stable state, whereas the presentation of shapes was stable for their duration. Such a difference enables us to test the extent to which duration is durable as a cue for grouping, over different modalities, and over dynamic versus static stimuli.

Examples of shape stimuli used in Experiment 2.

As with the speech, a familiarization sequence was created for each condition, comprising 150 shape triplets, with no shape triplet immediately repeated. Transitional probabilities between shapes within a triplet (transitional probability = 1) and between triplets (transitional probability = .33) were identical to those in Experiment 1. A blank screen occurred for 225ms between the presentation of every shape, as pilot studies demonstrated that without this the stimuli were uncomfortable to view. This meant that the stimuli were somewhat different than the speech streams in terms of continuity, however, again, this enables us to provide a stronger test of the robustness of the use of duration information for grouping stimuli across modalities. Streams in the initial- and final-item duration increase conditions lasted for 218 s, and streams in the no duration increase conditions lasted for 203 s.

For testing, we generated sequences that corresponded to words and part-words with the same structure as in Experiment 1: triplets that reliably occurred together during training were the “words,” and triplets that crossed boundaries between triplets were “part-words.” As with the speech, during testing all shapes were presented for the same duration (225 ms, with 225 ms blank screen interval): no duration cue was present at this stage. Pairs of sequences were separated by a 1000-ms pause.

Procedure

Participants were instructed as follows: “You will see sequences of shapes, and your task is to try to determine their structure.” Participants were assigned to a final-shape duration increase, a no duration increase, or an initial shape duration increase condition, and viewed the corresponding training sequences on a computer screen. At test, participants were asked: “Select which of two sequences best fits the structure of the sequences you just saw.” They then viewed the 16 forced-choice items, responding with a keyboard press as in Experiment 1. Training and testing was then repeated. All instructions were presented in the participants’ native language.

Results and Discussion

A repeated-measures ANOVA was performed on the data (proportion of correct responses), with cue condition (ISI, NI, FSI) and language group (English, Japanese) as between subjects factors, and test time (Test 1, Test 2) as a within subjects factor.

There was a significant effect of test time (Time 1: M = .597, SE = .016; Time 2: M = .705, SE = .021, F (1, 66) = 22.754, p < .001, η p 2 = .256). This may have been attributable to the effect of learning during the test, as “words” occurred more frequently than “part-words.” However, critically, all interactions involving test time were not significant and are not further reported (all p > .05), thus any learning during test did not affect performance in any of the duration conditions or language groups differentially.

There was no significant effect of language group, F < 1, indicating that English and Japanese listeners performed to a similar degree across the three conditions (ISI: English M = .595, SE = .019, Japanese M = .625, SE = .031; NI: English M = .635, SE = .042, Japanese M = .634, SE = .034; FSI: English M = .713, SE .044, Japanese M = .702, SE = .041).

There was a significant effect of cue condition, F (2, 66) = 3.820, p = .027, η p 2 = .104, and a linear contrast respecting the predicted order of cues (ISI, NI, FSI) was significant, F (2, 66) = 3.820, p = .027, η p 2 = .104 (see Figure 3 ). Dunnett’s post hoc t tests conducted to respect the hypothesized linear contrast order revealed that participants in the FSI condition ( M = .708, SE = .029) were more accurate than those in the NI condition ( M = .635, SE = .026), p = .047, and the ISI condition ( M = .610, SE = .018), p = .009. The NI and ISI conditions did not differ significantly, p = .379.

Mean shape-sequence identification score for participants in each cue condition, given for English and Japanese listeners. Error bars show ±1 SEM . See the online article for the color version of this figure.

Critically, there was no significant interaction between language group and cue condition, F (2, 66) = .209, p = .812, indicating that language background did not differentially affect use of duration information for learning the structure of visual sequences. The other interactions were not significant (all p > .05).

The results were very similar to those of the language stimuli in Experiment 1. The English listeners demonstrated the same benefit of element-final duration increase for grouping the statistically defined visual sequences. These results are consistent with previous studies of unstructured visual sequence processing, as shown by Peña et al. (2011) . In Peña et al.’s (2011) study, Italian participants, whose native language contains final-syllable duration increase ( Tyler & Cutler, 2009 ), demonstrated an iambic grouping preference for visual sequences varying in duration. Thus, our results corroborate these earlier findings, and demonstrate that they are generalizable to visual shape sequences with more complex statistical structure.

The results for the Japanese participants are, however, more difficult to reconcile with the hypothesis that grouping according to final element duration is a consequence of exposure to this structure in natural language. The remarkably similar use of final-element duration increase across modalities, and across listeners with different language backgrounds, suggests that the iambic preference operates independently of language exposure, and can be used to support statistical structure of sequences by participants with differing exposure to duration variation in their language experience.

Figures, tables, references, and supplementary files are best inspected in the licensed PDF or repository copy linked above.

Open Paper Agent