Open APA research · Launch collection CC BY 3.0

Mark my words: High frequency marker words impact early stages of language learning.

Frost RLA, Monaghan P, Christiansen MH.

Journal of experimental psychology. Learning, memory, and cognitionAmerican Psychological Association2019-01-17DOI 10.1037/xlm0000683

Abstract

High frequency words have been suggested to benefit both speech segmentation and grammatical categorization of the words around them. Despite utilizing similar information, these tasks are usually investigated separately in studies examining learning. We determined whether including high frequency words in continuous speech could support categorization when words are being segmented for the first time. We familiarized learners with continuous artificial speech comprising repetitions of target words, which were preceded by high-frequency marker words. Crucially, marker words distinguished targets into 2 distributionally defined categories. We measured learning with segmentation and categorization tests and compared performance against a control group that heard the artificial speech without these marker words (i.e., just the targets, with no cues for categorization). Participants segmented the target words from speech in both conditions, but critically when the marker words were present, they influenced acquisition of word-referent mappings in a subsequent transfer task, with participants demonstrating better early learning for mappings that were consistent (rather than inconsistent) with the distributional categories. We propose that high-frequency words may assist early grammatical categorization, while speech segmentation is still being learned. (PsycINFO Database Record (c) 2019 APA, all rights reserved).

Attribution and reuse record

Authors
Frost RLA, Monaghan P, Christiansen MH.
Original journal
Journal of experimental psychology. Learning, memory, and cognition
Publisher
American Psychological Association
Publication date
2019-01-17
DOI
10.1037/xlm0000683
License
CC BY 3.0
Open repository
Europe PMC · PMC6746567
Collection
School leadership launch collection

Presented by the Journal for School Superintendents under the license identified in the article’s open full-text record. The original authors and publisher do not endorse this journal or its agent.

Open full text

Read the scholarly record

Design

The experiment used a between-subjects design, with two conditions of training type: markers and no markers . These conditions varied the number of marker words present in the speech and either contained no marker words, or one marker word per category. Participants were randomly allocated to one of these conditions, with 24 participants receiving each type of training. All participants completed the same battery of tests after training; knowledge of the experimental language was tested immediately with tasks assessing speech segmentation and distributional categorization. A transfer task then put the familiarized language to use, to see whether participants’ distributional category knowledge for the target words shaped the way they used those targets as labels for actions and objects. For this task, participants were further subdivided into two groups for whom objects and actions were labeled in a way that was either consistent or inconsistent with the distributional categories (each N = 12). This subdivision was crossed with the markers/no markers conditions, such that half of the participants in each group received consistent labels, and the other half received inconsistent labels. Note that the no markers condition does not relate meaningfully to the consistent versus inconsistent distinction, but was included as an additional control to ensure that any effects observed for the markers condition were not due to biases in participants’ responses to individual items.

Sample size was designed with respect to our key experimental test of transfer from a speech segmentation task to a label-mapping task. We assessed the results from Graf Estes et al. (2007) and Frost et al. (2016) —studies in the literature which most closely resembled the current study design. Graf Estes et al. (2007) demonstrated a transfer effect of η p 2 = .16 in a two-way mixed ANOVA for an infant looking study to stimuli that either matched or mismatched object-labels, with 14 participants in each condition (conditions were word-transfer and nonword-transfer ). Post hoc power was .88, assuming zero correlation between matching and mismatching conditions (an intercorrelation would increase power further). Graf Estes et al. (2007) employed one within-subject and one between-subjects factor, whereas our design required two between-subjects factors. For a similar effect size, a priori power = .84 with 24 participants distributed across the marker word and no marker word conditions (with N = 12 in the consistent and inconsistent mapping conditions). For the pilot study by Frost et al. (2016) , three different between subject marker word conditions each had 24 participants, subdivided into groups receiving consistent and inconsistent mappings ( N = 12 for each subgroup). The size of the effect of marker word consistency on transfer to the early stages of word learning was η p 2 = .124, with power = .81.

In the current study we analyzed results using linear mixed effects models which can enhance power further ( Brysbaert & Stevens, 2018 ). As there were no similar studies using linear mixed effects analysis, we were unable to determine the required sample size for the study a priori. We thus report post hoc power for each of the measures we assessed below, generated using the simR package ( Green & MacLeod, 2016 ), as recommended by Brysbaert and Stevens (2018) .

This study received ethical approval from the Faculty of Science and Technology Research Ethics Committee at Lancaster University.

The experimental language and the stimuli and procedure for each of the tasks are outlined below.

Training

A continuous stream of synthetic speech was created using the Festival speech synthesizer ( Black et al., 1990 ) by concatenating target words and marker words (see Table 1 ). For the no markers control condition, the speech stream comprised target words only, and lasted approximately 280 seconds. For the markers condition, the speech stream comprised target words plus marker words, and lasted approximately 420 s. In both conditions, the eight target words were each presented 150 times with no immediate repetition, and speech was continuous, with no pauses between words. Speech streams had a 5-s fade in and out so that the onset and offset of speech could not be used as a cue to the word boundaries or language structure.

Segmentation test

To test segmentation, we created a two-alternative forced-choice task, which examined participants’ preferences for words versus nonwords. Nonwords were bisyllabic items that comprised the last syllable of one target word and the first syllable of another (e.g., muno formed from samu and noli ). We used nonwords (items which did not occur during familiarization) in order to make comparisons across the different conditions. For the no markers condition, particular transitions between target words were withheld from the speech stream, and nonwords were formed from the resulting syllable combinations of the omitted transitions (so for this group nonwords are comparable with part-words in classic instances of this paradigm). The same nonwords were used at test for the markers condition, and these did not occur in the familiarization speech for an additional reason: A marker word intervened between pairs of target words. Note that it would not have been possible to use part-words which spanned word boundaries as in Saffran, Aslin, and Newport’s (1996) studies of speech segmentation, because part-words did not occur in a comparable way across conditions: Part-words in the markers condition would comprise a fragment of a target word and a marker word, in the no markers condition, they would have to comprise fragments of two target words. In our task, preference for selecting words over nonwords would indicate that participants had successfully distinguished target words from competitor nonword syllable sequences.

Eight test pairs were constructed by matching each target word with a corresponding part-word (e.g., samu vs. muno ), and items in each test pair were separated by a 1-s pause. Test pairs were each presented twice, giving 16 test items in total. Items were presented in random order, and correct responses occurred an equal number of times in the first and second position within pairs.

Categorization test

To test abstract categorization (i.e., categorization based on the distributional information about the co-occurrence of markers and targets), we created an explicit similarity-judgment task that contained pairs of target words. Twelve test pairs contained items from the same category (as determined by the marker words that preceded them in speech), with six test pairs containing two Category A words, and six test pairs containing two Category B words. There were also 12 mixed test pairs, which contained one word from each category (so, one A word and one B word), giving 24 test pairs in total.

Transfer of category knowledge test

To test whether knowledge of distributional categories constrained participants’ learning of word-referent mappings, we created a cross-situational word-picture/action mapping task, which provided a grammatical category distinction (i.e., nouns and verbs) onto which the target words could map. On each trial, participants heard a sentence comprising two targets, then saw two visual scenes, with each scene containing a shape undertaking an action (with no duplication of shapes and actions on individual trials). Participants stated which of the two visual scenes the sentence described (see Monaghan et al., 2015 for a similar experimental design of cross-situational learning but without the preceding segmentation task, and see the Procedure section for more information about this task). There were four images of shapes, each printed in black on a grey background, taken from Fiser and Aslin (2002) . There were four actions that these shapes could perform: rotate, bounce, swing, and shake, selected from the series of actions used by Monaghan et al. (2015) .

Of the eight target words, four were paired with different shapes, and four were paired with different actions. Sentences were constructed to describe possible noun–verb combinations such that each sentence contained two target words, with one word referring to the shape, and one word referring to the action it was undertaking.

Critically, for half of participants, word-action/shape pairings were consistent with the distributionally defined categories heard during training, such that all A words appeared with shapes and all B words appeared with actions. For the remaining participants, pairings were inconsistent : two A words and two B words were paired with shapes, and two A words and two B words were paired with actions (see Figure 1 ).

Consistent versus inconsistent word-action/object mappings.

There were six versions of the language, presentation of which was counterbalanced across participants. Each version of the language used a different set of images for this task, which were selected at random from a set of eight novel shapes (taken from Fiser & Aslin, 2002 ). Particular objects and actions co-occurred an equal number of times, to prevent formation of associations between particular objects and actions, and to minimize unintentional co-occurrences between nouns and actions and between verbs and objects. Word-referent pairings differed across the six different versions of the language, to control for potential preferences for linking certain sounds to particular objects or actions.

Vocabulary test

Finally, we created a vocabulary test to assess exactly which word-object/action mappings participants had learnt. This task contained 16 two-alternative forced-choice trials, with two trials for each target word. Trials assessing learning of nouns (word-object mappings) contained static images of two objects: the target object and one other trained object that acted as a foil. For trials assessing learning of verbs (word-action mappings), we introduced a new shape to prevent participants’ responses from being influenced by knowledge of shapes. On verb trials, participants saw two scenes containing a new shape (presented alongside one another onscreen, as in the noun trials) with the shape performing a different action in each scene. On each trial, after 5 s (while the scenes were still onscreen) a target word was presented auditorily, and participants selected which of the two objects, or actions, the target was referring to. Pairing of shapes and actions was pseudorandomized such that each shape and action featured an equal number of times over the course of the task: twice as the correct referent (one to the left, and once to the right), and twice as the alternative (once to the left, and once to the right).

Procedure

Before hearing the familiarization speech, participants were instructed to pay attention to the language and think about the possible words it may contain. Participants were tested immediately after training. Testing was structured such that all participants received the tasks in the same order: Participants completed the segmentation test first, followed by the categorization test, then the cross-situational transfer test. Tasks were programmed using EPrime 2.0, with instructions appearing onscreen before each task began.

For the segmentation test, participants were instructed to listen to each test pair (a word and a nonword) then select which item best matched the language they had just heard, responding “1” for the first or “2” for the second sequence on a computer keyboard.

For the categorization test, participants were instructed to listen to each test pair (two words from the same/different categories), then rate how similar they thought the role of the items was in the familiarization stream. Participants were required to respond on a computer keyboard using a 6-point Likert-scale, with 1 = extremely different roles and 6 = very similar roles . If participants have formed categories based on the co-occurrence of target words and markers, then pairs of items taken from the same category should receive higher similarity ratings than mixed pairs (see Frost et al., 2016 , for a variant of this paradigm that demonstrates abstract category learning is possible).

For the cross-situational task, on each trial participants saw two scenes containing different objects performing different actions. After 5 s (while the scenes were still onscreen) participants heard a sentence comprising two target words. The sentence described one of the two scenes, with one word referring to the action, and one to the object. When the sentence had finished, participants were instructed to indicate via key press whether the sentence described the scene on the left or the right of the screen, by pressing “1” for the left and “2” for the right. The next trial began after participants had provided their response. An example trial is shown in Figure 2 .

An example trial on the cross-situational learning task, with two shapes performing unique actions, presented alongside a sentence that describes one of these pairs.

The probability of co-occurrence between a noun (noun i ) and its target object (object i ) was p (object i |noun i ) = 1, whereas co-occurrence between a noun and another distractor object was p (object j |noun i ) = 0.33, and co-occurrence probability between a noun and each action was p (noun i |action j ) = 0.25. Similarly, verb to target action co-occurrence probabilities were p (action i |verb i ) = 1, verb to other action probabilities in the distractor scene were p (action j |verb i ) = 0.33, and verb and object co-occurrence probabilities were p (object i |verb i ) = 0.25. Over the course of the task, we expected that learners would draw on these co-occurrence statistics in order to learn word-referent mappings.

There were six blocks, each containing eight learning trials. Within each block, each image and motion occurred four times—twice in the target scene and twice in the alternative (foil) scene. Each word occurred twice in each block. The left/right position of the referent and alternative scene was pseudorandomized such that each scene appeared once in each position.

To avoid providing additional cues for the role of words in the language, the presentation order for nouns and verbs was counterbalanced such that half of the sentences in a block followed a noun–verb order, and the other half followed a verb–noun order. Thus, this task used free word-order, such that grammatical categories of words were defined only by their prior co-occurrence with the marker words, and not in terms of the sentence position.

If prior category knowledge was influencing performance on this task, then participants should find it easier to use words from each of the categories consistently (i.e., all A words labeling objects) than inconsistently (i.e., some A words labeling objects, but some A words labeling actions). Thus, the key interaction of interest involves condition and consistency, which would reflect transfer of distributional category knowledge.

For the vocabulary test, on each trial participants saw the two scenes then heard a target word and selected via key-press whether the word referred to the scene on the left or the right of the screen (pressing “1” for left and “2” for right, as in the cross situational task). The next trial began after participants had responded. There were 16 trials in total, with two trials for each target word. The order of noun and verb trials was randomized, and the left/right position of the referent and foil alternative was balanced in the testing block.

Training and testing stimuli were presented at a comfortable volume, through closed-cup headphones. All participants were tested individually in an isolated booth, and the entire session lasted for approximately 30 min.

Results

We first report the results of the segmentation task, investigating the effect of marker words on participants’ ability to individuate words from the speech, relative to the no markers control group. We then present the results for the categorization test, which assesses whether participants encoded category information about the words on the basis of their co-occurrence with markers. Finally, we report the key analysis for the study, which is whether category information defined by the marker words can have an implicit effect on participants’ ability to use words as nouns and verbs in the word-object/action transfer task. For this test, we first report transfer effects seen during early stages of learning (consistent with learning effects observed in Frost et al., 2016 ), we then report the results across the whole task, followed by the measures of learning of nouns and verbs.

One-sample t tests were performed on the segmentation data (proportion correct responses) to compare performance with chance. Performance was significantly above chance for both no markers ( M = .740, SE = .042), t (23) = 5.658, p < .001; and markers ( M = .666, SE = .028), t (23) = 5.914, p < .001, indicating that participants in both conditions were able to identify individual words from the speech stream.

Generalized linear mixed effects analysis was performed on the data ( Baayen, Davidson, & Bates, 2008 ), modeling the probability (log odds) of response accuracy on the segmentation test considering variation across participants and materials. The model was built incrementally, and was initially fitted with the maximal random effects structure that was justified by the design, with random effects of subjects, particular test-pairs, and language version (to control for variation across the randomized assignments of phonemes to syllables). Random slopes were omitted if the model failed to converge with their inclusion ( Barr, Levy, Scheepers, & Tily, 2013 ). We then added condition (markers, no markers) as a fixed effect, and considered its effect on model fit with likelihood ratio test comparisons. There was no significant effect of condition (model fit improvement over the model containing random effects: χ 2 (1) = 2.850, p = .091, power = .46, 95% CI [.39, .53]), indicating participants in both conditions performed at a statistically similar level (difference estimate = −.459, SE = .27, z = −1.70, see Frost et al., 2016 for a similar observation in a pilot of this task). See Table 2 for a summary of the final model, and see the online supplemental materials for a visualization of the data for this task.

Figures, tables, references, and supplementary files are best inspected in the licensed PDF or repository copy linked above.

Open Paper Agent