Abstract
Disruptive behavior disorders (DBDs) are common in childhood and adolescence, with global estimates of 5.7%. While parenting practices are associated with DBDs, it is not clear whether these associations reflect causal effects or confounding. To strengthen causal inference, we meta-analyzed quasi-experimental evidence on the relationship between parenting practices and DBD symptoms. We conducted multilevel random-effects meta-analyses to pool results and assess evidence of heterogeneity and moderator analyses to further investigate potential sources of heterogeneity. We identified 45 studies that used data from 28 distinct cohorts (n = 38,591) and implemented seven different quasi-experimental methods. There was evidence of a causal effect of negative parenting practices on offspring DBD symptoms (Pearson's r = 0.13; 95% confidence interval, CI [0.09, 0.16]; 95% prediction interval, PI [-0.08, 0.35]; n = 30,677), but no effect of positive parenting practices (r = -0.06; 95% CI [-0.14, 0.02]; 95% PI [-0.39, 0.28]; n = 21,100). Moderator analyses indicated that the effect of negative parenting was consistent across offspring characteristics and maternal and paternal parenting but varied by type of quasi-experimental method, informant for the exposure and outcome, and study quality. The present study thus provides evidence of a small, harmful, causal effect of negative parenting practices on offspring DBDs. Effectively targeting such parenting practices could reduce the substantial societal burden of DBDs, with a potential 4% decrease in the global prevalence of DBD symptoms. This is equivalent to approximately 4.5 million school-aged children no longer meeting clinical thresholds for DBDs, which may reduce pressure on the criminal justice, health care, and social welfare sectors. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
Attribution and reuse record
- Authors
- Karwatowska L, Solmi F, Baldwin JR, Jaffee SR, Viding E, Pingault JB, De Stavola BL.
- Original journal
- Psychological bulletin
- Publisher
- American Psychological Association
- Publication date
- 2025-11-01
- DOI
- 10.1037/bul0000495
- License
- CC BY 4.0
- Open repository
- Europe PMC · PMC12720486
- Collection
- School leadership launch collection
Presented by the Journal for School Superintendents under the license identified in the article’s open full-text record. The original authors and publisher do not endorse this journal or its agent.
Open full text
Read the scholarly record
Study Selection
Studies meeting all of the following criteria were included in the review:
Included at least one clearly defined measure of parenting practices and at least one clearly defined measure of disruptive behavior.
Included a measure of parenting that was assessed either before or concurrently with the outcome.
Published in English, although the study could have been conducted in any country.
Used a quasi-experimental method (see definitions below).
Studies meeting any of the following criteria were excluded:
The study was a case report, clinical trial, systematic review, meta-analysis, or thesis.
The study used populations selected on physical health problems (e.g., cancer, seizures, surgery, low gestational age).
The study used populations selected on other diagnosed developmental disorders (e.g., language disorders, learning disorders, motor disorders, autism spectrum disorders) or mental health diagnoses (e.g., schizophrenia, depression, bipolar).
Positive and Negative Parenting Practices
We defined positive parenting practices as being warm, sensitive, or child-centered (e.g., use of praise or interest in offspring’s hobbies) and negative parenting practices as being harsh or insensitive (e.g., shouting, threatening behavior). We did not include physical discipline, abuse, or violence, as these do not represent “normative” parenting practices. We treated positive and negative parenting practices as separate constructs, as they are thought to have unique influences on offspring disruptive behavior ( Hipwell et al., 2008 ; Oliver et al., 2014 ; Pettit et al., 1997 ).
DBD Symptoms
We defined the outcome either by symptoms (e.g., conduct problems [CP], externalizing problems) or clinical diagnoses (e.g., CD, ODD, psychopathy, antisocial personality disorder) associated with disruptive behavior, which we refer to broadly as DBD symptoms. Further definitions are available in the Supplemental Table S5 .
Search Strategy
We searched Embase, APA PsycInfo, and MEDLINE for peer-reviewed studies written in English and published from January 1980 to April 2024. Search terms are reported in full in the Supplemental Table S6 and included terms relating to DBDs, parenting practices, and quasi-experimental methods. Two authors (LK and FS) independently screened the titles and abstracts of all articles retrieved from the searches. The full texts of all potentially eligible studies were also reviewed by two authors (LK and FS or BLDS).
Data Extraction
After the full-text screen, two authors (LK and JRB) independently extracted data from all eligible studies, including information on sample size, confounder adjustment, and effect sizes. The original study authors were contacted when this information was either missing or incomplete. When multiple effect sizes were available, the most conservative estimate (i.e., with the greatest degree of control for confounding) was extracted.
Risk of Bias
We adapted the Newcastle–Ottawa scale ( Wells et al., 2000 ) to include questions relevant to quasi-experimental studies. Additional/adapted questions included control for environmental and genetic confounders ( Supplemental Table S7 , Questions 5 and 6), whether the exposure and outcome were reported by different informants (Question 8), and whether the exposure and outcome were assessed longitudinally (Question 9). An overall score was derived by summing the scores across all items (highest possible score = 10), and the 33rd and 66th percentiles were used to categorize the studies into one of three categories used in the original Newcastle–Ottawa scale: “very high risk of bias” (score below 5.5), “high risk of bias” (score between 5.5 and 7), or “high quality” (score above 7). For studies that reported multiple effect estimates in different categories (e.g., high quality and high risk of bias), we gave the study an overall rating that corresponded to the highest category (e.g., high quality). One author (LK) coded study quality, and any questions were discussed with two members of the team (BLDS and J-BP).
Effect Size Transformation, Interpretation, and Significance
Most studies measured parenting practices and DBD symptoms on a continuous scale. If the effect parameters were not already standardized (i.e., reported as [Pearson’s correlations] r ), these were transformed into Pearson’s correlations using the formulae reported in Supplemental Table S8 . Therefore, the results from the meta-analyses represent the association between a 1 SD difference in a standardized parenting practices score and corresponding changes in a standardized offspring DBD score.
For negative parenting practices measures, a positive effect size ( r ) indicates that higher levels of negative parenting (e.g., more harsh or inconsistent discipline) are associated with more DBD symptoms. A negative effect size suggests that negative parenting is associated with lower levels of DBD symptoms. For positive parenting practices measures, a positive effect size indicates that higher levels of positive parenting (e.g., more warm and affectionate parenting practices) are associated with more DBD symptoms, while a negative effect size suggests an association with lower levels of DBD symptoms.
If standard errors of the reported parameters were not available, they were calculated using the sample sizes and reported p values.
Multilevel Random-Effects Model
All analyses were conducted in R (4.1.0) using the metafor (Version 4.3-7; Viechtbauer, 2010 ) package. As most studies ( k = 36; 80%) reported estimates for multiple measures of parenting practices and many studies ( k = 28; 62%) used data from the same data sources (i.e., the same cohort), we fitted three-level linear random-effects models ( Assink & Wibbelink, 2016 ) with the reported effect estimate nested within study nested within cohorts (see Supplemental Figure S1 ), which resulted in an overall “pooled” r .
To evaluate possible publication bias, we created funnel plots to check for asymmetry in the distribution of estimates according to their precision and conducted various additional analyses, including Egger’s test of heterogeneity ( Rodgers & Pustejovsky, 2021 ) and leave-one-out analyses to recalculate the Egger’s test when certain effect estimates were excluded ( Viechtbauer & Cheung, 2010 ).
We also examined potential heterogeneity using the Cochrane Q , I 2 , and τ 2 statistics. We interpreted an I 2 of more than 50% as an indication of moderate heterogeneity ( Higgins et al., 2019 ). To further investigate possible sources of heterogeneity, we conducted another set of leave-one-out analyses where we recalculated the Q , I 2 , and τ 2 statistics to see if statistical inferences changed when certain effect estimates were excluded.
Figures, tables, references, and supplementary files are best inspected in the licensed PDF or repository copy linked above.