Abstract
Remembering facilitates future remembering. This benefit of practicing by active retrieval, as compared to more passive relearning, is known as the testing effect and is one of the most robust findings in the memory literature. It has typically been assessed using verbal materials such as word pairs, sentences, or educational texts. We here investigate if memory for visual materials equally benefits from retrieval-mediated learning. Based on cognitive and neuroscientific theories, we hypothesize that testing effects will be limited to meaningful visual images that can be related to preexisting knowledge. In a series of four experiments, we systematically varied the type of material (meaningless "squiggle" shapes vs. meaningful object images) and the format of the test used to probe memory (a visually driven alternative forced-choice test vs. a remember/know recognition test). Within each experiment, we assessed the effects of practice type (retrieval or restudy) and the delay of the final test (immediate vs. 1 week) on the resulting practice benefits. Abstract shapes never showed a significant testing benefit, irrespective of test format. Meaningful object images did benefit from testing, particularly at long delays, and with a test format probing the recollective component of recognition memory. Together, our results indicate that retrieval can facilitate the recollection of visual images when they represent meaningful semantic units. This pattern of results is predicted by cognitive and neurobiologically motivated theories proposing that retrieval's benefits emerge through spreading activation in semantic networks, producing more easily accessible and longer-lasting memory traces. (PsycInfo Database Record (c) 2023 APA, all rights reserved).
Attribution and reuse record
- Authors
- Ferreira CS, Wimber M.
- Original journal
- Journal of experimental psychology. Learning, memory, and cognition
- Publisher
- American Psychological Association
- Publication date
- 2023-06-08
- DOI
- 10.1037/xlm0001248
- License
- CC BY 4.0
- Open repository
- Europe PMC · PMC10519161
- Collection
- School leadership launch collection
Presented by the Journal for School Superintendents under the license identified in the article’s open full-text record. The original authors and publisher do not endorse this journal or its agent.
Open full text
Read the scholarly record
Materials and Procedure
In this experiment, we used 40 (36 critical, four used for a familiarization task) black abstract shapes ( squiggles ; Figure 1C , upper yellow box), presented on white backgrounds. These stimuli were kindly shared by Groh-Bordin et al. (2006) and can be found at https://osf.io/6zf3t/ ( Ferreira & Wimber, 2021 ). Each squiggle was randomly paired with a word. Words were drawn from the MRC psycholinguistic database ( https://websites.psychology.uwa.edu.au/school/MRCDatabase/uwa_mrc.htm ) and had similar values of imageability ( M = 567.9, SD = 10.9), concreteness ( M = 583.7, SD = 30.7), and meaningfulness ( M = 447.4, SD = 31.1; all three scales measured in a range of 100–700).
The experiment consisted of four main stages: study, practice (retrieval or restudy, manipulated between subjects), immediate memory test, and delayed memory test ( Figure 1A and B ). The immediate test took place immediately after practice, whereas the delayed test took place after 7 days.
During study, participants saw a word in black font at the top of the screen, with a squiggle (4.8 × 6.4 cm) below for 7 s. Participants were instructed to link the word and the squiggle as well as they could, as they would be tested on the pairs later. In addition to memorizing the pairs, participants were asked to press a key to indicate whether they found it easy or hard to link the word and the squiggle together (left for easy, right for hard). Since these meaningless shapes are difficult to memorize (as revealed in a pilot study), each pair was repeated twice during the study phase. The order of stimulus presentation was randomized, but all the stimuli were presented once before repeating again in a new random order.
After study, a quick familiarization phase took place, to assess whether participants were paying attention to the pairs and to prepare them for what would be the format of the final test. This quick test also provided a break between study and practice, the longest phases of the experiment. During this familiarization task, participants saw four of the previously studied pairs. They were first presented with a word at the top of the screen and asked to think back to the squiggle associated with this word. After 4 s, a question mark appeared below the word, and participants were asked to indicate whether they thought they remembered the correct item (left arrow key) or not (right arrow key). Upon response, three different squiggles (all previously studied) appeared below the word, and participants were asked to pick the one that had been paired with that particular word by pressing one of the arrow keys on the keyboard (left for the leftmost stimulus on the screen, down for the middle stimulus, and right for the rightmost one). This 3-AFC screen disappeared upon participants’ response or after 4 s. The stimuli presented in the familiarization task were not shown again in the remaining parts of the experiment.
After familiarization, participants were informed that they would now have an opportunity to practice some of the pairs. Participants in the restudy condition (24/48) were told they would see some of the previously studied pairs and should use this reexposure as a chance to encode them again. The pairs were shown in the same way they had during study for 7.5 s, and participants had to indicate if they still found it easy or hard to link the pair, using the same response keys as in the study phase. Participants in the retrieval condition (24/48) were asked to actively bring the squiggles back to mind, upon being prompted with the word as a retrieval cue. The word appeared at the top of the screen with a question mark below for 5 s, during which participants were instructed to vividly bring the squiggle back to mind. The corresponding squiggle was then presented for 2.5 s to provide feedback. In both conditions, 24 of the 36 critical pairs were presented twice for practice. Stimuli were presented in random order, but all 24 items were presented once before repeating again in a new random order.
The remaining 12 pairs were not practiced and were used as baseline items to assess memory performance without further practice. This baseline measure was included to account for random variability in memory performance between participants and between the restudy and retrieval groups. Reducing such random noise is particularly relevant for the delayed test in our experiment, where differential forgetting rates are likely to increase variance in performance. Memory accuracy for nonpracticed baseline items was subtracted from accuracy for practiced items (see Statistical Analyses section), yielding a practice benefit for each participant that could then be compared between the two groups and between the immediate and delayed test.
The assignment of each word–image pair to practice or baseline was counterbalanced across participants. Note that both conditions (retrieval and restudy) were equated for overall practice time and were very similar, with the key difference that participants in the retrieval condition had to consciously bring the correct squiggle back to mind whereas participants in the restudy condition were presented with the complete pair.
Pairs were pseudorandomly allocated to the immediate or delayed test, so that half of all the pairs (12 practiced and six baseline) were tested immediately after practice, whereas the other half were tested 7 days later, all in random order. Other than that, the two tests were identical and followed the exact same procedure as the familiarization phase.
Statistical Analyses
The raw data, as well as the averaged data used in all the analyses throughout the manuscript, are available at https://osf.io/6zf3t/ . All statistical analyses used practice benefits (accuracy for practiced minus nonpracticed baseline items) as the dependent variable. Since testing effects are typically found after extended delays between practice and final test, we were particularly interested in retrieval benefits at a long delay. Accordingly, in all four experiments of this study, including the present one, we initially conducted one planned comparison, which was a one-tailed independent t test comparing the practice benefits between the retrieval and restudy group on the delayed test.
We were additionally interested in whether there was a shift from a restudy benefit in the immediate test to a retrieval benefit in the delayed test, as previously reported in the testing effect literature (e.g., Roediger & Karpicke, 2006b ). To assess this Practice Type × Delay interaction, we ran a 2 × 2 mixed analysis of variance (ANOVA) with practice benefit as the dependent variable, and factors practice type (retrieval vs. restudy; manipulated between subjects) and delay (immediate vs. delayed test; manipulated within subjects). Significant effects in this ANOVA were then further assessed in two-tailed post hoc t tests.
Results
The results of Experiment 1A are depicted in Figure 2A , showing practice effects for the squiggle images depending on the type of practice and delay (see Table 1 for results breaking down performance for practiced and baseline items). The one planned comparison of interest revealed no significant benefit of retrieval over restudy on the delayed memory test, t (46) = −2.44, p = .99. In fact, a significant effect in the opposite direction was found (see post hoc tests below).
Note . Colored rainclouds represent practice benefits for retrieved items, and gray rainclouds represent practice benefits for restudied items, in an immediate (left of the graphs) and delayed (right of the graphs) memory test. 3-AFC = three-alternative forced-choice. See the online article for the color version of this figure.
Results from the mixed ANOVA revealed no significant Practice Type × Delay interaction, F (1, 46) = 1.42, p = .239. There was no main effect of delay, F (1, 46) = 1.62, p = .210, but we did find a significant main effect of practice type, F (1, 46) = 4.08, p = .049, η p 2 = .08. Post hoc comparisons revealed that participants in the restudy group showed significantly larger practice benefits, across immediate and delayed test ( M = 0.12, SD = 0.23), than those in the retrieval group, M = 0.03, SD = 0.22; t (94) = −2.14, p = .035. This restudy advantage was statistically significant only on the delayed test, immediate test: t (46) = −.696, p = .490; delayed test: t (46) = −2.44, p = .019.
Discussion
In Experiment 1A, our main comparison of interest revealed no retrieval benefit for novel, meaningless shapes. In contrast, we found a reversal of the testing effect, with restudied items benefiting significantly more from practice than retrieved items at a longer delay.
There are several possibilities as to why a testing effect was not found here. One is that, as hypothesized, retrieval does not enhance memory for novel visual stimuli that have no preexisting representation in semantic memory. In fact, not only was no testing effect found for these meaningless squiggle images, but restudy seemed to benefit their long-term retention to a greater extent. This tendency for a restudy advantage was present at both delays, though only significant at the 1-week delay, with no interaction between practice type and delay. These findings are consistent with theories ascribing testing effects to the coactivation of semantically related information during retrieval ( Antony et al., 2017 ; Carpenter, 2009 , 2011 ; Pyc & Rawson, 2010 ; Sinclair & Barense, 2019 ). If the to-be-retrieved material has no existing semantic representation, spreading activation to similar information is not possible.
An alternative explanation is that retrieval practice for these meaningless shapes was simply too difficult. If participants in the retrieval group were largely unsuccessful at bringing back to mind and visualizing the correct items, this could potentially eliminate any practice benefits, and even make restudy the more advantageous rehearsal strategy. For example, previous work suggests that retrieval practice leads to substantial strengthening of only those items that are successfully recalled during practice. Restudy, by contrast, moderately strengthens all items uniformly ( Kornell et al., 2011 ; van den Broek et al., 2014 ). The bifurcation of the item strength distribution caused by retrieval practice can explain why restudy sometimes outperforms retrieval on immediate tests: the moderate strength of restudy items is sufficient to support these items’ recall when little forgetting has happened. However, retrieval will outperform restudy on delayed tests, where forgetting has pushed most restudy items below the accessibility threshold, while the subset of items that were successfully retrieval practiced will remain accessible ( Rowland & DeLosh, 2015 ). Note that this account does not provide a mechanistic explanation for the different processes underling restudy and retrieval practice. It does, however, predict that retrieval benefits are limited to items successfully retrieved, or corrected by feedback ( Rowland & DeLosh, 2015 ). Such failure to retrieve, however, is unlikely to have caused our pattern of results. First, feedback was provided for 2.5 s on each retrieval trial, exposing participants to the correct item even if they had not been able to retrieve it. Secondly, baseline performance levels were comparable between the retrieval and restudy group on the immediate final test, and even numerically higher in the retrieval group on the delayed test (see Table 1 ), speaking against an effect of low retrieval practice success.
Finally, a third possibility that could account for our results is that the final memory test used in this experiment was not sensitive to retrieval benefits. As Chan and McDermott (2007) have pointed out, testing effects are evident only when the final test specifically encourages controlled retrieval of the studied items. Across four experiments, these authors and others ( Pu & Tse, 2014 ; Verkoeijen et al., 2011 ) demonstrated that recognition tests that rely heavily on familiarity often fail to detect any differences between retrieval and restudy conditions, with differences becoming evident, however, when final memory tests specifically probe recollection.
While previous research suggests that AFC tests depend on recollection (e.g., Cook et al., 2005 ; Kroll et al., 2002 ; Khoe et al., 2000 ), especially when using familiar lures as in the present design ( Mayes et al., 2002 ; Migo et al., 2009 ), others have argued that discrimination in these tests can be achieved on the basis of familiarity (e.g., Bastin & Van der Linden, 2003 ).
To minimize the contribution of familiarity and isolate the recollective component of memory retention on the final test, we conducted the same experiment again, now using a remember/know associative recognition procedure as the final memory test instead of the 3-AFC. Associative recognition tests are thought to depend strongly on recollection ( Hockley & Consoli, 1999 ), especially when participants are asked to judge the oldness of stimuli that are all familiar but presented in a rearranged fashion (e.g., reshuffled study pairs). In this case, familiarity is less useful in supporting recognition (since all of the items are equally familiar; Yonelinas et al., 2010 ), and the rejection of rearranged pairs requires recollection ( Castel & Craik, 2003 ). However, associative recognition has also been shown to be subject to low-level perceptual influences ( Goshen-Gottstein & Moscovitch, 1995 ). To account for this, we added remember/know judgments to our final test to specifically isolate the recollection component of the recognition process. In this procedure, a “remember” response is thought to reflect recollection processes, whereas “know” responses should be based on familiarity ( Gardiner, 1988 ; Migo et al., 2012 ; Tulving, 1985 ). In Experiment 1B (and also Experiment 2B), we thus counted only correct “remember” responses as successfully retrieved, allowing us to isolate the benefits of retrieval and restudy practice specifically on recollection-based memory (for analyses including “know” responses, see pages 1 and 2 in the online supplemental materials ).
Materials and Procedure
The main difference between Experiments 1A and 1B was the final test, where an associative recognition test including remember/know judgments was used instead of the 3-AFC test. For the associative recognition test, additional squiggle images were selected as lure images to be shown on repaired trials. Of the 40 squiggles shown at study, 20 were knotted (their lines crossed at one point of the drawing) whereas the other 20 were simple squiggles (i.e., not knotted—no lines crossed; see upper yellow box in Figure 1C ). This feature was used to select perceptually similar lures for the final test (see below). Other than that, the study phase was the same as in Experiment 1A.
The familiarization phase served again as a preparation for the final tests. As before, participants saw the cue word for 4 s and were asked to think back to the associated squiggle . Then, a question mark appeared below, and participants were asked to report by button press whether they remembered the correct item or not. They were then presented with an item below the word. Participants had to indicate if the item had originally been presented with that same word or not (see below). The item was on screen until response or up to a maximum of 4 s.
The squiggle shown, together with a word, in the familiarization task and the final tests could be (a) exactly the same that had been studied in the first phase of the experiment (original pairs), (b) a squiggle that had never been seen before, but was perceptually similar to the studied one (perceptual lures) or (c) a squiggle and a word that had both been previously seen in the experiment, but had not been paired together (episodic lures; see Figure 1C ). Knotted squiggles served as perceptual lures for other knotted squiggles as did simple for simple. The participants’ task was to decide if a given pairing was old (intact) or new (repaired). They were made aware of the different types of pairs and were instructed to respond with “old” only to the original pairs and “new” to the two types of repaired probes (i.e., perceptual and episodic lures). Moreover, if the item was old, they were asked to indicate whether they remembered (that is, they distinctively remembered seeing the item and the word paired together during the study phase of the experiment) or knew it (had the feeling they had seen the pair before, without precise recollection). Participants pressed the left arrow on the keyboard for “old-remember,” the down arrow for “old-know” and the right arrow for “new.” These prompts were shown at the bottom of the screen, below the squiggle , in the left, middle, and right positions, respectively ( Figure 1B ). In the familiarization phase, four original pairs, one episodic lure, and one perceptual lure were shown in random order.
After the familiarization phase, participants performed the practice phase twice. The retrieval and restudy conditions were identical to Experiment 1A. The only difference was that we asked participants in the retrieval condition for a subjective memory response before the probe squiggle appeared on the screen by pressing the left button if they thought they remembered the correct item and pressing the right button if they did not remember the item. This button press was included to provide us with an indication of memory success, even though subjective, which was not available for Experiment 1A. In pages 3–5, the online supplemental materials report these subjective judgments. As in the previous experiment, all items were presented once in random order before repeating again, in a new random order. Assignment of each pair to practice or baseline, and of each squiggle to target or lure, was counterbalanced across participants.
The final tests (immediate and delayed) followed the same procedure as the familiarization phase. Pairs were pseudorandomly chosen within each participant’s learning set to be tested immediately or after a week, so that at each test stage, 18 original pairs, 18 episodic lures, and 18 perceptual lures were tested. Of these, 12 were previously practiced items (or lures of practiced items) and six were baseline items (or lures of baseline items). Half were knotted and half were simple squiggles .
In this experiment, we used two types of lure items: perceptual and episodic. If retrieval strengthens the meaningful aspects of a memory ( Ferreira et al., 2019 ; Lifanov et al., 2021 ), this might come at the cost of perceptual detail ( Lifanov et al., 2021 ) and increase false alarms to perceptually similar lures ( Lee et al., 2017 ). For instance, repeatedly retrieving the image of a set of keys (see example in Figure 1 ) will presumably activate and strengthen the existing concept “key” ( Antony et al., 2017 ) at the expense of finer perceptual details of the specific set of keys that had been studied (see Lee et al., 2017 ). Accordingly, we additionally hypothesized that in our experiments using the associative recognition final memory test (Experiments 1B and 2B), retrieval (compared to restudy) would specifically increase perceptual, but not episodic, false alarms.
Finding perceptual lures in Experiment 2B, where concrete objects were presented as stimuli (see below), was relatively straightforward. For example, we selected a set of keys that is perceptually similar to the target set of keys, but not the same (see Figure 1C ). For the present experiment using abstract squiggle stimuli, the selection of lure images is more difficult. To keep Experiments 1B and 2B as coherent as possible, we still aimed to approximate the perceptual lure manipulation here. Stimuli in the present experiment included knotted and simple squiggles (see Figure 1C for examples), and we used these two categories to draw perceptual lures from the same category as the target squiggle ; that is, if the target was a knotted squiggle , so was the perceptual lure, while simple squiggles were used as lures for simple target squiggles .
Statistical Analyses
Like in Experiment 1A, our dependent variable was practice effects, calculated here as the proportion of original pairs that participants correctly recollected (old-remember responses) after practice compared to no practice. As mentioned earlier, our aim was to isolate the effects of testing on recollection, and we therefore only counted old-remember responses as correctly retrieved to obtain a maximally pure measure of recollection ( Gardiner, 1988 ; Tulving, 1985 ). Results using old-know and all old responses collapsed are reported in pages 1 and 2 of the online supplemental materials . Briefly, these analyses showed a significant effect of delay (where practice benefits were more pronounced in the delayed test) but no other significant effects. These results should be interpreted with caution, given the low number of “know” responses.
Consistent with the previous experiment, we first conducted a targeted one-tailed t test comparing the practice benefits (old-remember responses to practiced—baseline items) in the retrieval and restudy groups on the final delayed test. We then ran a 2 × 2 mixed ANOVA with factors practice type (retrieval vs. restudy; manipulated between subjects) and delay (immediate vs. delayed test; manipulated within subjects). Significant results from the ANOVA were further assessed in two-tailed post hoc comparisons.
Finally, we analyzed the proportion of false alarms to perceptual and episodic lures. Because participants rarely gave a “remember” response to new pairings, “remember” and “know” false alarms were collapsed for these analyses. We analyzed old responses to lures of practiced pairs minus old responses to lures of baseline items to parallel all other analyses. We were particularly interested if retrieval, compared to restudy, would increase the proportion of perceptual false alarms ( Lee et al., 2017 ).
Results
Practice effects in Experiment 1B are depicted in Figure 2B , and Table 1 shows mean accuracies separately for practiced and baseline items. The main planned comparison showed no significant retrieval advantage for squiggles in the delayed test, t (45) = −.232, p = .59. Moreover, there was no significant Practice × Delay interaction, F (1, 45) = .913, p = .344, nor a main effect of practice, F (1, 45) = 1.40, p = .243, or delay, F (1, 45) = 3.09, p = .085. No further post hoc tests were thus conducted on the practice benefits.
Analyzing false alarm rates, we found no differences in perceptual false alarms (old responses to perceptual lures of practiced minus of baseline items) between the retrieval and restudy groups on the delayed test, t (45) = .926, p = .82. Additionally, the 2 × 2 ANOVA indicated no significant interaction between practice type and delay, F (1, 45) = .008, p = .93, nor a main effect of practice, F (1, 45) = 2.34, p = .133. There was, however a main effect of delay, F (1, 45) = 10.67, p = .002, with participants showing a larger practice-related increase in perceptual false alarms on the delayed test compared to the immediate test when collapsing across both groups, M imm = −.097, SD imm = 0.18; M del = .034, SD del = 0.18; t (46) = −3.304, p = .002. No significant effects were found when using false alarms to episodic lures (old responses to episodic lures of practiced minus baseline items) as the dependent variable for any of the planned analyses, t test on delayed test: t (45) = 1.17, p = .88; Delay × Practice interaction: F (1, 45) = 1.97, p = .17; main effect of practice: F (1, 45) = .083, p = .74; main effect of delay: F (1, 45) = .017, p = .90.
Discussion
Experiment 1B again provided no evidence for a testing effect when using meaningless squiggle images. In contrast with Experiment 1A, we did not find a significant reversal of the testing effect in this study either, although numerically, practice benefits where still higher for the restudy condition (see Table 1 ). Together with Experiment 1A, this pattern of results suggests that retrieval does not enhance memory for images that have no preexisting semantic representation, irrespective of the final test format. The finding is congruent with predominant theories of the testing effect suggesting that spreading activation in semantic networks plays a central role in producing retrieval’s benefits on long-term retention ( Antony et al., 2017 ; Carpenter, 2011 ; Ritvo et al., 2019 ; Sinclair & Barense, 2019 ): novel, meaningless materials can be assumed to preclude such spread of activation due to a lack of a preexisting associative network, resulting in no testing effect. Instead, restudy can be beneficial in such situations, allowing for additional exposure to the novel materials.
It should be noted, however, that the null findings from Experiment 1B in themselves do not provide strong evidence for or against any theory. If the lack of preexisting knowledge explains the absence of a testing effect in our first two experiments, we should expect a change in direction of the practice benefits when the target images depict meaningful objects, with a clear testing effect emerging on the delayed test.
We thus conducted two further experiments, replacing the abstract squiggles with concrete nameable objects. In Experiment 2A, we used a 3-AFC final memory test, whereas in Experiment 2B, participants’ memory was probed with an associative recognition test including remember/know judgments, mirroring Experiments 1A and 1B, respectively. We hypothesized that a retrieval-induced enhancement should be evident particularly in Experiment 2B, where the memory test specifically probes recollection, replicating previous work ( Chan & McDermott, 2007 ; Pu & Tse, 2014 ; Verkoeijen et al., 2011 ).
Figures, tables, references, and supplementary files are best inspected in the licensed PDF or repository copy linked above.