Open APA research · Launch collection CC BY 3.0

Statistical learning for speech segmentation: Age-related changes and underlying mechanisms.

Palmer SD, Hutson J, Mattys SL.

Psychology and agingAmerican Psychological Association2018-09-24DOI 10.1037/pag0000292

Abstract

Statistical learning (SL) is a powerful learning mechanism that supports word segmentation and language acquisition in infants and young adults. However, little is known about how this ability changes over the life span and interacts with age-related cognitive decline. The aims of this study were to: (a) examine the effect of aging on speech segmentation by SL, and (b) explore core mechanisms underlying SL. Across four testing sessions, young, middle-aged, and older adults were exposed to continuous speech streams at two different speech rates, both with and without cognitive load. Learning was assessed using a two-alterative forced-choice task in which words from the stream were pitted against either part-words, which occurred across word boundaries in the stream, or nonwords, which never appeared in the stream. Participants also completed a battery of cognitive tests assessing working memory and executive functions. The results showed that speech segmentation by SL was remarkably resilient to aging, although age effects were visible in the more challenging conditions, namely, when words had to be discriminated from part-words, which required the formation of detailed phonological representations, and when SL was performed under cognitive load. Moreover, an analysis of the cognitive test data indicated that performance against part-words was predicted mostly by memory updating, whereas performance against nonwords was predicted mostly by working memory storage capacity. Taken together, the data show that SL relies on a combination of implicit and explicit skills, and that age effects on SL are likely to be linked to an age-related selective decline in memory updating. (PsycINFO Database Record (c) 2018 APA, all rights reserved).

Attribution and reuse record

Authors
Palmer SD, Hutson J, Mattys SL.
Original journal
Psychology and aging
Publisher
American Psychological Association
Publication date
2018-09-24
DOI
10.1037/pag0000292
License
CC BY 3.0
Open repository
Europe PMC · PMC6233520
Collection
School leadership launch collection

Presented by the Journal for School Superintendents under the license identified in the article’s open full-text record. The original authors and publisher do not endorse this journal or its agent.

Open full text

Read the scholarly record

Neuropsychological Testing

This study was conducted as part of a larger aging project in which all participants completed a battery of neuropsychological tests. Within the context of the present experiment, the variables of interest were those relating to working memory and executive function, including forward and backward digit spans, working memory updating, and inhibition (measured using a Stroop task). Additionally, processing speed was measured using the Digit Symbol-Coding and Symbol Search subtests of the Wechsler Adult Intelligence Scale ( Tulsky, Zhu, & Ledbetter, 1997 ). Full descriptions of the working memory tasks are provided in the online supplemental materials . The Mini Mental State Exam (MMSE; Folstein, Folstein, & McHugh, 1975 ) was administered to the older age group to screen for signs of abnormal cognitive decline. All participants scored above the standard cutoff of 24 points on the MMSE.

SL recognition test

The recognition test for the SL task was a two-alternative forced-choice task composed of 24 pairs. The 24 pairs included eight word versus part-word trials (WD-PW), eight word versus nonword trials (WD-NW), and eight part-word versus nonword trials (PW-NW). For WD-PW trials, a word from the familiarization phase was paired with a part-word, that is, a word from the other stream of the same language. Part-words were named as such because they occurred across word boundaries in the familiarization stream. For WD-NW trials, a word from the familiarization phase was paired with a nonword. Nonwords were syllable strings created using syllables heard in the familiarization stream, but concatenated in a scrambled order so that any two consecutive syllables never occurred in that order in the familiarization stream. For PW-NW trials, a part-word was paired with a nonword. The PW-NW trials were included as filler trials to ensure that part-words, nonwords, and words appeared equally often during the test phase. Within each trial type, each syllable string appeared twice, each time paired with a different string. Across all trials, syllable strings were only paired with strings of the same length (i.e., a quadrisyllable was never paired with a trisyllable). The order in which the word pairs were presented was randomized between participants.

Cognitive load task

For the cognitive load conditions, a 2-back task was created. This consisted of rapid serial presentation of visual stimuli displayed concurrently with the speech stream. The stimuli were 86 line drawings taken from Kroll and Potter (1984) , which were all novel, meaningless, and nonnameable shapes. Half of the shapes were rotated 30° to the left and the other half were rotated 30° to the right. The 2-back repetition trials always involved a change in orientation: if the first shape was rotated to the right, the repetition was rotated to the left, and vice versa. Each shape was displayed for 750 ms. Shapes were separated by 250 ms of blank screen. The pace was the same regardless of the speech rate. It was chosen to avoid synchrony between the onset of the visual stimuli and the onset of words or syllables in the stream. The total number of shapes displayed depended on the duration of the speech stream. The normal rate included 481 shapes with 78 2-back pairs, whereas the slow rate included 864 shapes with 140 2-back pairs.

Design and Procedure

The experiment followed a 2 × 2 design with cognitive load and speech rate as within-subjects variables. Each participant completed the following four conditions: (a) normal rate with no cognitive load, (b) normal rate with cognitive load, (c) slow rate with no cognitive load, and (d) slow rate with cognitive load. In each condition, participants were tested using a different language. Each of the four conditions was completed during one of four individual testing sessions. The session timings were arranged in accordance with the participant’s availability, with the constraint that there should be at least 48 hr and no more than 2 weeks between two consecutive sessions. The order in which the conditions were completed was counterbalanced between participants. Each session lasted approximately one hour and consisted of an SL task and a subset of neuropsychological tests (detailed above). Participants always completed the SL task before performing the neuropsychological tests. All participants completed the audiometric test at the start of Session 1. Hearing thresholds were tested for 250, 500, 1000, 2000, 4000, and 8000 Hz.

The experiment took place in a sound-attenuated booth. Stimuli were played over headphones. Several steps were taken to ensure that the sound level was appropriate for each individual participant. First, each participant performed an audiometric test before taking part in the experiment. Participants were then played nonsense trisyllabic sequences over the headphones and asked whether the sound level was comfortable and whether they could hear the syllable sequences clearly. Sound level was adjusted accordingly. Finally, participants performed a syllable intelligibility test in which they heard 16 syllables (four from each of the four languages). After each syllable, participants were asked to repeat what they heard. If a syllable was repeated incorrectly, the experimenter corrected it and replayed the syllable. One participant in the older age group was eliminated and later replaced due to poor performance in the syllable test.

Participants were told that they would hear an artificial language and that they should try to discover what the words of the language were. To ensure that participants understood the instructions and knew what to expect, they were always played a sample stream for 30 s prior to starting the familiarization phase. The sample stream was played at the same rate as the real subsequent familiarization stream. It included a combination of syllables from each of the four languages, but did not include words from any of them. In the cognitive load conditions, participants also practiced the cognitive load task while listening to the sample stream. During the first session, participants were shown a sample trial of the test phase after hearing the sample stream so that they understood the format of the test. Before hearing the experimental stream, participants were informed that the stream would be made up of artificial words that were three and four syllables in length. However, they were not told how many words were in the stream or the proportion of trisyllables and quadrisyllables. Since our design was within subjects, these procedures were employed to minimize the impact of “novelty” during the first session, and ensure that the participant’s expectations at the start of the first session were as similar as possible to those at the start of the later sessions.

In the cognitive load (2-back) conditions, participants were instructed to press the space bar every time they saw a shape that was the same as the shape that appeared two trials before ( Figure 1 ). Participants were informed that the shape might have a different orientation on its second presentation.

Recognition test

For each trial in the recognition task, the two strings of a pair were presented both visually and auditorily. The two strings were separated by a 500-ms silent interval and were played at the same rate as the speech stream heard during familiarization. As the first string was played, its orthographic transcription appeared on the left-hand side of the computer screen for 1,500 ms and disappeared before the second string was played. The orthographic transcription of the second string then appeared on the right-hand side of the screen as the second string played, and stayed for 1,500 ms. Then, both transcriptions appeared again, simultaneously, and remained on the screen until the participant responded. Supplementing the auditory strings with visual support was intended to prevent memory decay between the first and second strings, which could have excessively affected the older participants. Participants were asked to indicate which syllable string had occurred more frequently in the artificial language using the left or right shift key for the first or second syllable string in each pair, respectively. The next pair was presented 1,000 ms after the participant’s response.

Performance in the 2-Back Task

Hit rates, false alarm rates, and d ′ scores for the 2-back task are shown in Table 1 . 1 The hit rate was calculated as the number of correct responses to 2-back repetitions divided by the total number of 2-back repetitions in the stream. The false alarm rate was calculated as the total number of incorrect responses to nonrepeated stimuli divided by the total number of nonrepeated stimuli. Since d ′ scores are based on hit and false alarm rates aggregated over a large set of individual responses, the d ′ scores were analyzed through a two-way mixed ANOVA with stream rate (normal, slow) as the within-subjects variable and age group (young, middle-aged, older) as the between-subjects variable. This analysis showed that stream rate did not affect discrimination performance, F (1, 91) = 2.73, MSE = .38, p = .10, but there was a marginal difference in discrimination performance according to age group, F (2, 91) = 2.58, MSE = .70, p = .08.

Neuropsychological Measures

Mean performance in the neuropsychological tests for each age group is shown in Table 2 (see Footnote 1 ). Average hearing thresholds are also included. Table 3 shows the correlations between performance on each of the neuropsychological tests and performance in the SL task, both overall, and split by load and trial type.

In order to determine how well individual differences in working memory, executive function, and processing speed predicted SL performance when controlling for age and hearing, hierarchical multiple regression analyses were performed. Given our hypothesis that performance on WD-PW trials may be more dependent on higher level executive function than performance on WD-NW trials, we ran separate regression analyses for WD-PW and WD-NW trials. In both analyses, age and hearing were entered as predictors in the first block of the regression. Forward digit span, backward digit span, working memory updating, Stroop, and processing speed were entered in a second block. In order to minimize any outlier-related bias in the regression models, we eliminated extreme scores from the data by removing participants whose performance was more than 2 SD above or below the mean on any of the neuropsychological tests. This resulted in the loss of four participants: one from the middle-aged group and three from the older group.

In the WD-PW condition ( Table 4 ), Block 1 was not significant, R 2 = .05, p = .09, but Block 2 was, R 2 = .19, p = .02, with working memory and processing speed increasing the explained variance from 5% to 19%, Δ R 2 = .13, p = .03. In Block 2, working memory updating was the only significant predictor, β = .25, p = .03, showing that better working memory updating was associated with better performance. The other variables did not have a significant unique contribution.

The results of the WD-NW regression analysis are shown in Table 5 . As in the previous analysis, hearing and age did not predict performance in Block 1, R 2 = .01, p = .78. Again, the second block was significant, R 2 = .16, p = .04, with working memory and processing speed increasing the explained variance from 1% to 16%, Δ R 2 = .16, p = .01. This time, however, forward digit span was the only significant predictor of performance, β = .30, p = .04, showing that a larger working memory storage capacity was related to better performance on WD-NW trials. The other variables did not have a significant unique contribution.

In sum, age, hearing, and processing speed did not predict SL performance in either regression analysis. Working memory resources contributed to SL, but, critically, the type of resources involved seemed to depend on the nature of the computation performed. For the WD-NW trials, SL performance was related to general working memory storage capacity, with forward digit span accounting for 30% of the variance on these trials. In contrast, for the WD-PW trials, performance was related to working memory updating, which accounted for 26% of the variance. This shows that performance on WD-PW trials is more dependent on higher level executive function, and specifically, the ability to actively update the content of working memory. Interestingly, performance on WD-PW trials was not predicted by forward digit span, or by our other measures of executive function (backward digit span or inhibition [Stroop]).

Discussion

The purpose of the current study was to (a) investigate the effect of aging on speech segmentation by SL, and (b) gain further insight into the mechanisms that underlie this ability. In relation to our first question, the results indicate that, in general, SL is remarkably resilient to age-related decline. The performance of the middle-aged adults did not differ from that of young adults in any of the conditions tested; in fact, it was numerically slightly better. Perhaps more surprisingly, across all trials, SL in the no-load condition was almost equivalent in the young and older adults. At first glance, this suggests that SL is akin to other forms of implicit learning, to the extent that it is characterized by relative stability across the life span, compared with explicit forms of learning (e.g., paired-associate learning), which tend to show more marked age-related decline (e.g., Naveh-Benjamin, 2000 ; Naveh-Benjamin, Guez, Kilb, & Reedy, 2004 ; Naveh-Benjamin, Hussain, Guez, & Bar-On, 2003 ). However, when the data were examined more closely, evidence of age-related decrement in SL was visible in the older group under some circumstances. Specifically, age effects emerged in the more challenging conditions. A detailed examination of performance across these conditions provides insight into the mechanisms that underlie SL.

Despite the generally comparable performance across age groups, the older adults tended to perform less well than the other groups on the more difficult WD-PW trials. In contrast, performance on the WD-NW trials was largely unaffected by age. Since the part-words used as foils on WD-PW trials were present as such in the stream, it can be argued that WD-PW trials constitute a stronger test of SL than WD-NW trials, in which the nonword foils were never heard before. Indeed, succeeding on WD-PW trials is contingent on acquiring distinct and precise representations of the words in the stream, whereas succeeding on WD-NW trials only requires a general gist of familiar-sounding sequences. Therefore, although older adults appeared unimpaired when performance was considered across all trial types, age deficits were visible when more sensitive measures of SL were used. Likewise, the older adults were the only group who did not perform above chance on the WD-PW trials in the cognitive load conditions. The young and middle-aged groups both continued to learn under cognitive load, albeit to a lesser extent than in the no-load conditions. These findings suggest that older adults have more difficulty using transitional probabilities to form accurate word representations than younger adults, and that this age-related deficiency is exacerbated under cognitive load.

The fact that older adults failed to learn under cognitive load is consistent with the hypothesis that working memory resources contribute to SL. Since working memory typically shows some age-related decline (e.g., Bopp & Verhaeghen, 2005 ; De Beni & Palladino, 2004 ; Fiore et al., 2012 ; Van der Linden et al., 1994 ), older adults are likely to have fewer resources available to cope with the dual-task demand. Palmer and Mattys (2016) previously suggested that working memory updating may be particularly important for SL. This hypothesis was supported by analysis of the neuropsychological test data which revealed that performance on WD-PW trials was predicted by working memory updating; participants with stronger memory updating scores tended to perform better on WD-PW trials. Since the older adults performed worse on the working memory updating task than the other age groups, this could explain, at least in part, why they had more difficulty with the WD-PW trials.

There are at least two ways in which working memory updating might benefit SL performance. One possibility is that updating benefits SL indirectly. Here, participants who are better at updating are necessarily better at coping with the cognitive load task since the 2-back is essentially a memory updating task. However, it seems unlikely that this is the only way in which updating benefits SL. As shown in Table 3 , the correlation between working memory updating and SL performance was comparable in the load and no-load conditions for WD-PW trials. It was also the case that updating scores predicted SL performance only on WD-PW trials. They did not significantly predict SL performance on WD-NW trials, for which short-term memory storage capacity, as indexed by forward digit span, was relatively more important. This dissociation is important because it enables us to link performance on WD-PW trials specifically to the updating function of working memory, rather than to general working memory capacity. It seems likely, therefore, that working memory updating also benefits SL directly by supporting the acquisition of specific word-form knowledge required for WD-PW discrimination. This provides more concrete evidence for Palmer and Mattys’ (2016) findings which indicate that executive resources are recruited during SL. They previously suggested that working memory updating may assist SL by removing and replacing erroneous syllable grouping from working memory, leading to the more accurate representations required to distinguish words from other familiar-sounding sequences.

Performance on the easier WD-NW trials was less affected by age. This finding is consistent with the results of our regression analysis which revealed that working memory storage capacity, as measured by forward digit span, rather than working memory updating, was the strongest predictor of performance on these trials—note that the older adults performed no worse than the young adults on measures of working memory capacity. It seems likely that performance on WD-NW trials relied more on a general feeling of familiarity with test sequences than on computation of transitional probability. Interestingly, within the recognition memory literature, it has frequently been reported that older adults sh

Figures, tables, references, and supplementary files are best inspected in the licensed PDF or repository copy linked above.

Open Paper Agent