Abstract
Research suggests that sleep deprivation both before and after encoding has a detrimental effect on memory for newly learned material. However, there is as yet no quantitative analyses of the size of these effects. We conducted two meta-analyses of studies published between 1970 and 2020 that investigated effects of total, acute sleep deprivation on memory (i.e., at least one full night of sleep deprivation): one for deprivation occurring before learning and one for deprivation occurring after learning. The impact of sleep deprivation after learning on memory was associated with Hedges' g = 0.277, 95% CI [0.177, 0.377]. Whether testing took place immediately after deprivation or after recovery sleep moderated the effect, with significantly larger effects observed in immediate tests. Procedural memory tasks also showed significantly larger effects than declarative memory tasks. The impact of sleep deprivation before learning was associated with Hedges' g = 0.621, 95% CI [0.473, 0.769]. Egger's tests for funnel plot asymmetry suggested significant publication bias in both meta-analyses. Statistical power was very low in most of the analyzed studies. Highly powered, preregistered replications are needed to estimate the underlying effect sizes more precisely. (PsycInfo Database Record (c) 2022 APA, all rights reserved).
Attribution and reuse record
- Authors
- Newbury CR, Crowley R, Rastle K, Tamminen J.
- Original journal
- Psychological bulletin
- Publisher
- American Psychological Association
- Publication date
- 2021-11-01
- DOI
- 10.1037/bul0000348
- License
- CC BY 3.0
- Open repository
- Europe PMC · PMC8893218
- Collection
- School leadership launch collection
Presented by the Journal for School Superintendents under the license identified in the article’s open full-text record. The original authors and publisher do not endorse this journal or its agent.
Open full text
Read the scholarly record
Impact of Total Sleep Deprivation After Learning and Potential Moderators of the Effect
The active systems consolidation theory suggests that sleep after learning strengthens new memories (e.g., Klinzing et al., 2019 ; Kumaran et al., 2016 ; McClelland et al., 1995 ). Information learned during wakefulness is initially encoded rapidly in the hippocampus, where memories are stored separately from existing memory stores. Repeated reactivation of these new memories, primarily during slow-wave sleep (SWS), supports the strengthening of memory representations and leads to the integration of these memories in the neocortex. Such neocortical representations are less liable to disruption and form interrelated semantic networks with existing memories, yielding memory representations that allow abstraction, generalization, and discovery of statistical patterns across discrete memories ( Lewis & Durrant, 2011 ; Lewis et al., 2018 ; Stickgold & Walker, 2013 ). Notably, since the mechanisms outlined in this theory relate only to hippocampal-dependent memory consolidation ( Inostroza & Born, 2013 ), these mechanisms may primarily support the consolidation of declarative (or explicit) memory.
Effects of sleep on nondeclarative memory have also been observed; for example, sleep enhances motor skills such as finger-tapping sequence learning ( Korman et al., 2007 ; Walker et al., 2002 ; see King et al., 2017 for a review). However, this beneficial effect of sleep on procedural memories may be evident only when learning is intentional (explicit memory), rather than unintentional (implicit memory). Robertson et al. (2004) found that awareness of learning a finger-tapping task led to a sleep benefit, whereas improvements in implicit learning performance, when participants had little awareness of the task, were similar regardless of whether the retention interval contained sleep or wakefulness. Thus, there are suggestions that the mechanisms involved in the consolidation of hippocampal-dependent declarative memories may also facilitate the consolidation of some procedural tasks that rely on explicit learning and thus show some hippocampal-dependency ( Schönauer et al., 2014 ; Walker et al., 2005 ).
However, this latter theory does not take into account observations that some procedural tasks that do not rely on explicit learning or an intact hippocampus still show superior performance after sleep. Stickgold et al. (2000) found that a period of sleep after learning a visual discrimination task benefited later performance, and Schönauer et al. (2015) found that improvements in performance on a mirror-tracing task were only observed after a period of offline consolidation. Recent studies in both animals ( Sawangjit et al., 2018 ) and humans ( Schapiro et al., 2019 ) have suggested that the hippocampus may be involved in the sleep-dependent consolidation of memories that do not rely on the hippocampus during encoding. For example, Schapiro et al. (2019) trained amnesic patients with hippocampal damage on the motor sequence task, a classic procedural memory task typically considered to be non-hippocampus-dependent. The patients were able to learn the task equally well compared to matched controls, suggesting that the hippocampus is not required for learning of the task. However, while the controls showed the expected overnight consolidation benefit, the patients did not, leading the authors to conclude that the hippocampus may be involved in the consolidation of procedural tasks that do not require it for initial learning.
As reviewed above, theories of memory consolidation predict that depriving participants of sleep after learning should impair memory for the information encoded before sleep deprivation, relative to control conditions where participants are allowed to sleep normally after learning. In our first meta-analysis, we analyze the current research into both declarative and procedural memories to estimate the size of this sleep deprivation after learning effect. We focus on the literature using manipulations of total sleep deprivation, as this is the strongest and most direct manipulation to test theories of sleep-associated memory consolidation. In doing so, we exclude from our analyses studies of sleep restriction. In standard sleep restriction studies, participants are allowed to sleep for a shorter duration than they would do otherwise, and the manipulation typically continues for multiple nights. This chronic sleep restriction can have a detrimental impact on learning and memory (e.g., Cousins et al., 2018 ) although not always (e.g., Voderholzer et al., 2011 ). In selective sleep restriction studies, participants are only deprived of the first half of the night, rich in SWS, or the second half of the night, rich in rapid eye movement (REM) sleep (e.g., Plihal & Born, 1997 ). These studies can reveal important information about the precise brain mechanisms underlying benefits of sleep on memory. There are two theoretically motivated reasons for excluding both types of sleep restriction from our meta-analyses. Standard sleep restriction studies still allow participants to sleep for several hours each night, and it may be sufficient for sleep-associated memory consolidation to occur. Thus, these manipulations do not provide a strong test of the relevant theories, and their inclusion might lead us to underestimate the effect size associated with sleep deprivation. Selective sleep restriction studies on the other hand are designed to address the more fine-grained question of which specific sleep stages are the most beneficial for memory. Different theories make different predictions in this regard: For example, the active systems consolidation theory emphasizes the role of SWS ( Klinzing et al., 2019 ); the theory of Lewis et al. (2018) emphasizes REM (at least for learning involving creative thought); and the sequential theory of Giuditta (2014) proposes that an interaction between SWS and REM is key. However, all current theories make the common prediction that sleep should benefit m
Impact of Total Sleep Deprivation Before Learning and Potential Moderators of the Effect
Recent research has also proposed a role for sleep before learning. According to the synaptic homeostasis hypothesis ( Cirelli & Tononi, in press ; Tononi & Cirelli, 2003 , 2012 ), learning occurs during wake when a neuron detects a statistical regularity in its input and begins to fire in response to this regular input. In other words, successful learning requires neurons to be able to fire selectively in response to statistically regular patterns observed in the environment. To do so, strength of the synapses carrying these inputs must be increased. However, the neuron now faces the plasticity-selectivity dilemma. As an increasing number of input lines become strengthened, a larger range of input patterns can make the neuron fire, reducing the neuron’s ability to fire selectively. This loss of selectivity corresponds to reduced ability to encode new information. During sleep, the brain spontaneously activates both new information encoded during previous wake and information encoded in the past. Over the course of this activation, those synapses that are activated most strongly and consistently during wake survive, while at the same time, those synapses that were less activated are weakened. This weakening occurs primarily during the transitions between intracellular up and down states experienced during SWS. This competitive down-selection of weaker synapses restores memory encoding ability.
The restorative function of sleep is supported by evidence showing decreased episodic learning ability across a 6-hr retention interval in which participants remained awake, whereas encoding capacity was restored after a daytime nap ( Mander et al., 2011 ). Further, neuroimaging evidence suggests sleep deprivation prior to learning is associated with disrupted encoding-related functional activity in the bilateral posterior hippocampus ( Yoo et al., 2007 ; for similar findings, see Alberca-Reina et al., 2015 ; Drummond & Brown, 2001 ; Van Der Werf et al., 2009 ). Thus, sleep deprivation before learning may be detrimental specifically for the encoding of hippocampal-dependent declarative memories. Our second meta-analysis seeks to estimate the effect size associated with this impairment. As only two studies have looked at the impact of sleep deprivation on procedural learning when it occurs after deprivation, we were not able to assess the potential moderating effect of declarative versus procedural memory. We were also not able to use emotionality as a moderator here due to the low number of relevant studies. Yet, there are several studies that have used recall tasks and recognition memory tasks, so we were able to evaluate the moderating effects of recall versus recognition, as in the first meta-analysis.
The Present Meta-Analyses
Despite the breadth of evidence for an effect of total sleep deprivation both before and after learning on memory performance, there is no comprehensive review and analysis of the strength of the effect of sleep deprivation on long-term memory. Previous reviews and meta-analyses investigating a role of sleep deprivation have focused on tasks that are likely more susceptible to fatigue. Pilcher and Huffcutt (1996) conducted a meta-analytic review of the effect of sleep deprivation on cognitive and motor task performance in 19 primary studies and found that sleep deprivation had a significant impact on performance. Still, this meta-analysis does not address long-term memory performance. Similarly, Lim and Dinges (2010) found an effect of sleep deprivation on a range of cognitive tasks including attention, working memory, and short-term memory, though the size of the effect varied depending on the task (e.g., a nonsignificant effect on reasoning accuracy, but a large effect on attention). Finally, Harrison and Horne (2000b) in a review found that sleep deprivation impacted decision-making ability. The tasks studied in these reviews are often repetitive and monotonous (e.g., the Psychomotor Vigilance Task; Dinges & Powell, 1985 , the go/no go paradigm, and tests of serial addition), and they tend to probe lower-level functions that are particularly susceptible to fatigue, such as reaction times and processing speed. Therefore, the conditions of these studies are arguably better suited to finding adverse effects of sleep deprivation on performance than studies looking at higher-level learning and long-term memory. Thus, a review of the effects of sleep deprivation on such high-level, long-term memory is required.
Taking a meta-analytic approach will permit not only a quantitative assessment of the size of the main effect of sleep deprivation and its moderators but also an investigation of methodological quality within this literature including the statistical power of studies proposing to find a sleep deprivation effect. Variety in sample selection and methodological designs used within this literature raises the possibility of variations in methodological quality. Such variations could lead to biases in the meta-analysis by overestimating or underestimating the effect size ( Higgins et al., 2011 ). Thus, we developed a checklist to assess multiple aspects of methodological quality, including questions specifically relevant to the assessment of sleep effects (e.g., excluding participants with sleep disorders), questions on study design (e.g., within-subject vs. between-subjects design, and random allocation to conditions), and questions on data analysis (e.g., preregistration and a priori power analyses). We used similar questions to other meta-analyses in the sleep literature ( Lim & Dinges, 2010 ; Lo, Groeger, et al., 2016 ; Schäfer et al., 2020 ) and examined methodological quality as a continuous moderator in the analysis. The full methodological quality checklist is provided in Supplemental Appendix A .
It has been suggested that psychological science more broadly is currently suffering from a replication crisis due to low power, publication bias, selection biases, and analysis errors ( Nosek et al., 2015 ). Low power limits potential to detect genuine effects but also results in Type I errors and exaggerated effect sizes ( Ioannidis, 2005 ; Pollard & Richardson, 1987 ). Szucs and Ioannidis (2017) conducted an analysis of almost 4,000 cognitive neuroscience and psychology papers and found that the overall mean power to detect small, medium, and large effects was 17%, 49%, and 71%, respectively, with even lower power in the subfield of cognitive neuroscience. Given the convention that power to detect an effect size should be at least 80% ( Di Stefano, 2003 ), it is clear that a large number of studies within psychology are underpowered (see Button et al., 2013 ; for further evidence of low statistical power within neuroscience). In the sleep literature, sample sizes tend to be low, potentially due to the resource intensity of conducting these experiments. Thus, we investigated whether the low power seen more broadly in psychological science and neuroscience is also evident in the sleep deprivation literature. For each individual effect size entered into the meta-analysis, we calculated the study’s power (defined as power to detect our meta-analytic effect size) and investigated whether there is an association between a study’s power and the effect size observed in the study.
Inclusion/Exclusion Criteria
To select relevant studies, we applied the following inclusion/exclusion criteria.
Participants had to be healthy adults aged 18 years and older.
Studies must have included, as a primary independent variable, a manipulation of sleep deprivation that was a minimum of one night of total sleep deprivation with an appropriate sleep control condition consisting of one normal night of sleep. Residency studies (studies conducted in a medical setting) were excluded due to the lack of control over whether total sleep deprivation occurred (sleep deprivation is often reported despite naps having occurred on shift; e.g., Bartel et al., 2004 ; Guyette et al., 2013 ). Additionally, studies using sleep restriction protocols, which involve multiple nights of limited sleep duration rather than one or more nights of no sleep, were excluded because the neural and cognitive effects of sleep restriction may differ from those caused by total sleep deprivation ( Banks & Dinges, 2007 ; Lowe et al., 2017 ).
Studies must have included, as a primary dependent variable, at least one measure of learning or long-term memory where the task was described in sufficient detail to ascertain which cognitive function it assessed.
For the meta-analysis investigating sleep deprivation after learning, the cognitive task must have had a single encoding phase and a retrieval phase(s) that were temporally separated by either a period of sleep deprivation or an equivalent period of sleep. For the meta-analysis investigating sleep deprivation before learning, the single encoding phase and the retrieval phase(s) must have been temporally separated by a retention interval that had a minimum duration of at least 1 min, rather than being part of the same task session. The reason for this criterion is that our meta-analyses aimed to investigate effects of sleep deprivation on learning and long-term memory. The inclusion of studies with temporally indistinct encoding and retrieval phases would have included short-term and working memory tasks that form a separate body of literature ( Lim & Dinges, 2010 ), the analysis of which was beyond the scope of these meta-analyses.
In cases in which studies assessed the effects of other interventions (e.g., caffeine; Kilpeläinen et al., 2010 ) in ameliorating sleep deprivation effects, studies were included only if data could be obtained from the control sleep deprivation and control sleep groups. This criterion was included because the goal of the current meta-analyses was to assess effects of sleep deprivation in the absence of alertness-promoting strategies.
Studies must have reported sufficient statistical detail to calculate effect sizes (means, SD , F , and t ). When statistical details were not reported in the text, we either contacted corresponding authors to request relevant data or we extracted the data needed from published figures in the article using WebPlotDigitizer ( Rohatgi, 2019 ).
Methodological Quality
Through our survey of the literature, it became clear that sleep deprivation studies differ considerably in various aspects of methodological rigor (e.g., lack of control over adherence to sleep manipulations in the sleep deprivation and sleep control groups; Fischer et al., 2002 vs. complete control; Chatburn et al., 2017 ). For this reason, we assessed the methodological quality of each study entered into our meta-analyses and included this in our moderator analyses.
To assess methodological quality, we developed a 22-item checklist based on criteria for standard sleep deprivation experiments (e.g., preexperimental sleep monitored using actigraphy and exclusion of sleep disorders) and more general experimental psychology experiments (e.g., a priori power analysis and study design). For each item on the checklist, studies were scored with either a zero or a one according to whether they satisfied the criterion. To transform the total methodological quality score for each study into a risk of bias that reflects a rank of all the studies on a common scale, we normalized the total scores by dividing each study’s total methodological quality score by the maximum total methodological quality score that was achieved among all studies ( Stone et al., 2020 ). Lower values imply lower ranked studies (minimum score of 0) and higher values imply higher ranked studies (maximum score of 1) relative to the best study. The full methodological quality checklist can be found in Supplemental Appendix A .
Given that the checklist items form a multidimensional scale, the items were clustered according to the Downs and Black’s (1998) instrument for assessing methodological quality, which assesses five types of bias: “reporting,” “internal validity—bias,” “internal validity—confounding,” “power,” and “external validity.” The “reporting” cluster determines whether sufficient information is provided to make an unbiased assessment of study findings. In our methodological quality checklist, the items in this cluster referred to the reporting of exclusion criteria for participant characteristics (e.g., “Did the study exclude participants with a history of sleep disorders?”). The “internal validity—bias” cluster assesses whether biases were present in the intervention or outcome measure that would favor one experimental group [e.g., “Was interference for the sleep deprivation group low (nondemanding activities given and monitored in the lab)?”]. The “internal validity—confounding” cluster assesses whether biases were present in the selection and allocation of participants (e.g., “For within-group studies, was the order of deprivation and control conditions counterbalanced?”). The “power” cluster assesses whether a study used a priori power analyses to avoid Type II errors (e.g., “Did the study report an a priori power analysis with power set at 80% or higher and an α at .05 or lower?”). The Downs and Black’s (1998) “external validity” cluster determines the extent to which findings can be generalized to the population from which a sample was taken (e.g., “Were the staff, places, and facilities where the patients were treated, representative of the treatment the majority of patients receive?”; Downs & Black, 1998 , p. 383). Since the items in this cluster were designed for clinical intervention studies with nontypical populations, we dropped the external validity cluster from our checklist. In line with Cochrane Collaboration recommendations ( Higgins et al., 2011 ), the four clusters in our methodological quality checklist (hereon referred to as reporting, bias, confounding, and power) were then included in moderator analyses. The percentage of studies that passed on each item of the quality checklist for both Meta-Analysis 1 and Meta-Analysis 2 can be found in Supplemental Appendix D . Total methodological quality scores for each study, as well as the item-level ratings, can be found at osf.io/5gjvs/ .
Effect Size Calculation
Information on study means, standard deviation, and effect sizes for each item, as well as formulas used to calculate effect sizes, can be found at osf.io/5gjvs/ . We report the standardized mean difference in task performance between a sleep deprivation and sleep control group, with positive values indicating that sleep deprivation influenced learning and memory such that performance was significantly worsened compared to a sleep control group. For studies with independent samples (between-subjects designs), we computed Cohen’s d s based on the means and variance reported in each study for the sleep and sleep deprivation group. For within-subject designs, in which participants took part in both the sleep deprivation and sleep control conditions, we calculated Cohen’s d av , as recommended by Lakens (2013) .
Heterogeneity
To investigate whether moderating variables may influence the size of the effect of sleep deprivation, we examined heterogeneity within the data set using the Q test ( Cochran, 1954 ). The Q test indicates whether there is heterogeneity within the data set and is calculated by the weighted sum of the squared deviations of individual study effect estimates and the overall effect across studies. Significant heterogeneity suggests that some of the variance within the data set may not be due to random sampling error, and thus moderating variables may influence the effect. Since we were interested in both the within-study and between-study variance, we ran two separate one-sided log-likelihood-ratio tests. As such, the fit of the overall multilevel model was compared to the fit of a model with only within-study variance and to a model with only between-study variance. This allowed us to determine whether it was necessary to account for both within- and between-study variances within our model.
Assink and Wibbelink (2016) suggest that such log-likelihood ratio tests may be subject to the issues of statistical power when the data set comprises a small number of effect sizes. Low statistical power may lead to nonsignificant effects of heterogeneity when in fact there is variance within or between studies. To account for this, it is recommended to also calculate the I 2 statistic, which indicates the percentage of variation across studies that is due to heterogeneity and that which is due to random sampling error ( Higgins & Thompson, 2002 ). Hunter and Schmidt (2004) suggest the 75% rule, such that if less than 75% of overall variance is attributed to sampling error, then moderating variables on the overall effect size should still be examined. Using the formula of Cheung (2014) , we calculated the percentage of variance that can be attributed to each level of our model.
However, although I 2 reports the proportion of variation in observed effect sizes, it does not provide us with absolute values that tell us the variance in true effects ( Borenstein et al., 2017 ). Thus, as recommended by Borenstein et al. (2011) , we report the τ 2 , which provide an estimate of the true effect size, and we report prediction intervals, which indicate that 95% of the time, effect sizes will fall within the range of those prediction intervals.
Publication Bias
To assess publication bias, we first examined a contour enhanced funnel plot. Funnel plots show each effect size plotted against its standard error, with contour lines corresponding to different levels of statistical significance. If studies are missing almost exclusively from the white area of nonsignificance, then there may be publication bias. If studies are missing from areas of statistical significance, the bias is likely due to other causes such as poor methodological quality, true heterogeneity, chance, or the bias may be artifactual ( Johnson, 2021 ; Sterne et al., 2011 ). We also conducted a variation of Egger’s regression test for funnel plot asymmetry ( Egger et al., 1997 ) that can be conducted with multilevel meta-analyses.
Figures, tables, references, and supplementary files are best inspected in the licensed PDF or repository copy linked above.