Abstract
In an attempt to better characterize the complexity of difficulties observed within developing populations, numerous data-driven techniques have been applied to large mixed data sets. However, many have failed to incorporate the core role of developmental time in these approaches, that is, the typical course of change in behavioral features that occurs over childhood to adolescence. In this study, we utilized manifold projections alongside a gradient-boosting model on data collected from the Millennium Cohort Study to unpack the central role of developmental time in how behavioral difficulties transition between the ages of 5, 11, and 17. Our analysis highlights numerous observations: (a) Girls develop relatively greater internalized behavioral problems during adolescence; (b) in the case of a chaotic home environment, co-occurring internalizing and externalizing difficulties tend to persist during childhood; (c) peer problems were the most likely to persist over the whole 12-year period (especially in the presence of early maternal depression and poor family relationships); and (d) there were two pathways with distinct risk factors leading to antisocial behaviors in adolescence-an early-childhood onset pathway and later adolescent onset pathway. Our findings provide evidence that investigations of child and adolescent difficulties must be open to the possibility of multiple subgroups and variability in trajectory over time. We further highlight the crucial role of family and social support and school experience-related factors in predicting children's outcomes. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
Attribution and reuse record
- Authors
- Benhamou E, Akarca D, Bathelt J, Fletcher-Watson S, Astle DE.
- Original journal
- Developmental psychology
- Publisher
- American Psychological Association
- Publication date
- 2025-02-27
- DOI
- 10.1037/dev0001874
- License
- CC BY 4.0
- Open repository
- Europe PMC · PMC12243395
- Collection
- School leadership launch collection
Presented by the Journal for School Superintendents under the license identified in the article’s open full-text record. The original authors and publisher do not endorse this journal or its agent.
Open full text
Read the scholarly record
Measures
Table 2 lists all tests administered at each sweep.
Behavioral profiles were obtained from Goodman’s SDQ ( Goodman, 1997 ) completed by the main caregiver at ages 5, 11, and 17 years old. The SDQ contains five subscales: Emotional Symptoms, Conduct Problems, Hyperactivity, Peer Problems, and Prosocial Behaviors (reversed to Antisocial Sehavior in this study). Example items from each subscale include: “Often unhappy, down-hearted or tearful” (Emotional Symptoms); “Often has temper tantrums or hot tempers” (Conduct Problems); “Restless, overactive, cannot stay still for long” (Hyperactivity); “Generally liked by other children” (Peer Problems); and “Shares readily with other children (treats, toys, pencils etc.)” (Prosocial Behaviors). Scores ranging from 0 to 10 for each of the five subscales are obtained by adding up parental responses from each of its five items, which are rated on a 3-point Likert-type scale (0 = not true , 1 = somewhat true , or 2 = certainly true ). After reverse scoring the original Prosocial subscale, higher scores reflect greater difficulties. The validated three-band categorization was used to define two groups of children: (a) presenting with difficulties in any of the five SDQ subscales (part of the “borderline” or “abnormal” bandings), referred to as the “elevated” group, having relatively high difficulty scores, and (b) those who did not have any parent-reported difficulties across any scale, referred to as the “non-elevated” group.
Numerous studies have demonstrated the convergent and discriminant validity of the SDQ ( van Widenfelt et al., 2003 ; Vogels et al., 2011 ). Specifically, when compared to the gold standard Child Behavior Checklist or Youth Self-Report scale(s), the SDQ scales are more strongly associated with its conceptually similar Child Behavior Checklist/Youth Self-Report scale(s) than with conceptually different Child Behavior Checklist/Youth Self-Report scales (for more details, see Vugteveen et al., 2021 ). The SDQ has also been validated across multiple clinical groups. For instance, for ADHD, the Hyperactivity/Inattention scale has a sensitivity of 91% and a specificity of 90%, and for autism, three scales combine to provide a sensitivity of 79% and specificity of 93% ( Russell et al., 2013 ). However, it is important to recognize that it is not valid for all conditions. In a community sample of almost 8,000 5- to 15-year-olds, the SDQ identified individuals with a psychiatric diagnosis with a specificity of 95% and a sensitivity of 63%. However, the questionnaire was markedly better at detecting individuals with conduct, hyperactivity, depressive, and some anxiety disorders, identifying over 70% of them. But this detection rate drops to below 50% for those with specific phobias, separation anxiety, and eating disorders ( Goodman et al., 2000 ).
Cognitive Performance
A detailed guide on the cognitive measures used within the MCS study is available ( Moulton et al., 2020 ). Specifically, direct assessment of cognitive performance and mental health were obtained at the three time points. The earlier sweeps of the MCS consisted predominantly of the British Ability Scale II (BAS II) measures ( Elliott et al., 1996 ).
The cognitive assessment for MCS3 (Sweep 3; 5 years old) consisted of:
BAS II, Naming Vocabulary. This is a verbal test which assessed the spoken vocabulary of the child. Test items consist of colored pictures of objects shown one at a time, and the cohort member responds verbally as to what the object is. The test comprises of 36 items (pictures of objectives) in total and has been used to predict delayed language development from 3 years ( Law et al., 2012 ).
BAS II, Pattern Construction. This is a nonverbal, spatial, problem-solving test. For each item, a pattern was presented to the cohort member, and the cohort member was asked to replicate the pattern using flat form squats or solid plastic cubes with black and yellow patterns on each side. The test comprises of 23 items and has been used in modeling early-years school attainment ( Sullivan et al., 2013 ).
BAS II, Picture Similarity. This is a nonverbal pictorial reasoning test. For each item, the cohort member was shown a row of four pictures or designs, and the cohort member placed a fifth card below the stimulus picture it best matched. The test comprises of 33 items and has been investigated in its moderating role on family and neighborhood risk ( Flouri et al., 2015 ).
The cognitive assessment for MCS5 (Sweep 5; 11 years old) consisted of:
BAS II, Verbal Similarities ( Elliott et al., 1996 ). This is a verbal test that assessed the child’s verbal reasoning using verbal concepts. The child was given three stimulus words and asked to name the class to which all the examples belong. The test comprises of 35 items.
CANTAB Cambridge Gambling Task ( Atkinson, 2015 ). This assessed the child’s decision-making and risk-taking behavior. The child would be asked to bet on the number of points to bet on before the decision is made. Intelligence has been associated with risk adjustment and quality of decision making derived from this task ( Flouri et al., 2019 ).
The cognitive assessment for MCS7 (Sweep 7; 17 years old) consisted of:
Granada Learning Assessment Number Analogies ( Granada Learning Assessment, 2024 ). This assessed the young person’s arithmetic knowledge and reasoning with numbers.
Mental Health
Mental health metrics were also different across age sweeps. At age 5, we used two subscales of the Child Social Behavior Questionnaire completed by the main caregiver measuring the child’s independence and self-regulating behavior and the child’s emotional dysregulation ( Hartman et al., 2006 ). Items included in these subscales include “Likes to work things out for self” and “Easily frustrated,” respectively.
At ages 11 and 17, a shortened five-item Rosenberg Self-Esteem scale was also asked of cohort members ( Rosenberg, 1965 ). Eleven-year-old children’s answer to the question “How do you feel about your life as a whole?” was used as a proxy for overall happiness level (1 = completely happy to 7 = not happy at all ).
A range of different mental health measures were added at age 17, including the self-completed Neuroticism subscale of the Big Five personality traits (often listed under the acronym OCEAN: Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism), assessing the young person’s emotional reactiveness and vulnerability to stress; the self-completed six-item Kessler scale, which measures non-specific psychological distress ( Kessler et al., 2003 ), including items such as “During the last 30 days, about how often did you feel so depressed that nothing could cheer you up?”; the self-completed, short, seven-item Young Person Warwick–Edinburgh Mental Well-being Scale ( Stewart-Brown et al., 2009 ) which provides a single summary score indicating overall well-being, including items such as “I’ve been feeling optimistic about the future”; the answer to whether cohort members have ever been diagnosed with depression and/or serious anxiety; and the answer to whether they have ever attempted to end their life.
Longitudinal and Cross-Sectional Risk Analysis
Having used the measures above to compare groups after clustering, our second analysis drew on a large number of factors, looking at longitudinal and cross-sectional risk factors for behavioral transitions between the ages of 5, 11, and 17 years old. We selected these on the basis of the extant literature. We included factors from all MCS sweeps, including MCS1 (Sweep 1; 9 months old) and MCS2 (Sweep 2; 3 years old), preceding or concurrent to the transition “source” sweep. In other words, these factors could originate from any time up and until the “take off” point for the behavioral transition.
Factors were grouped into nine domains: peri- and postnatal factors (birth weight, type of delivery, feeding problems; Bhutta et al., 2002 ; McCormick et al., 1990 ; Serati et al., 2017 ; Wiles et al., 2006 ), household factors (e.g., housing tenure, neighborhood safety, average family income; Adjei et al., 2022 ; Clair, 2019 , p. 20; Miller et al., 2021 ; Minh et al., 2017 ), child physical health (e.g., BMI, sleep; Griffiths et al., 2011 ; Heron et al., 2013 ; Holley et al., 2011 ; Turnbull et al., 2013 ), child well-being (e.g., reluctance to go to school, self-esteem; Donnellan et al., 2005 ; Egger et al., 2003 ), child cognition (e.g., language development, mathematical abilities; Baird et al., 2022 ; Quistberg & Mueller, 2020 ; Thompson et al., 2019 ), caregiver health and mental health (e.g., Kessler scale, use of recreational drugs, alcohol consumption; Adjei et al., 2022 ; Maselko et al., 2015 ; Meadows et al., 2007 ; Schepman et al., 2011 ), caregiver relations (e.g., marital status, satisfaction with partner; Giannakopoulos et al., 2009 ), child relations (e.g., time spent outside with friends, time spent on social media; Brunborg & Burdzovic Andreas, 2019 ; Van Roy et al., 2010 ), and caregiver–child relationship (e.g., style of parenting, closeness to child; Bloomfield & Kendall, 2012 ; Coldwell et al., 2006 ). See Supplemental Table 3 for the full list of factors.
Analysis Procedure
Our analysis pipeline was designed first to clean the data and partition it into those with elevated and non-elevated difficulty scores on the SDQ (note that elevated here includes children with borderline or elevated scores on at least one scale). Then, with the elevated group, we conducted a dimension reduction approach called UMAP, before clustering, identifying significant transitions across development, and then establishing significant risk factors for these transitions ( Figure 1 ; Table 3 ). Readers may wonder why we partitioned the data to focus on those with elevated scores. This is necessary because otherwise the dimensional space is dominated by individuals with scores within the normal range and not optimized to capture behavioral difficulties. In turn, the data present lower dimensionality because it is dominated by the extreme differences in overall severity across the whole sample and is then very difficult to cluster meaningfully.
Note . XGBoost = gradient-boosting algorithm for classification; SDQ = Strength and Difficulties Questionnaire; UMAP = uniform manifold approximation and projection; KNN = k -nearest neighbor; max_depth = maximum depth (of our tree); min_child_weight = minimum child weight (minimum sum of instance weight needed for a child of our tree); reg_alpha = regularisation alpha (L1 regularisation term on weights). See the online article for the color version of this figure.
a In the UMAP projection of SDQ, data from five dimensions (5D) to two dimensions (2D) reduce the size of the data but in a way that we do not lose out an essential information underlying the data. Subsequent figures present the 2D plot of each subject’s reduced 2D data.
The first step of our analysis was to identify subgroups of children within the elevated sample at each time point by using a clustering method on z -scored SDQ derived values. Before clustering, we first excluded any participants with missing SDQ data or missing age or gender information. We then partitioned our children into the elevated and non-elevated difficulty groups. We then removed univariate (>3 SD s from the median) and multivariate outliers (Mahalanobis distance >α quantile of the chi-square distribution with 5 df ) separately for each group. The total number of children excluded at each time point is indicated in the consort diagrams in Supplemental Figures 1–3 (MCS 3: n = 702 or 4.7%, MCS 5: n = 60 or 0.5%, MCS 7: n = 61 or 0.7%, for univariate or multivariate outliers or because of missing information).
To avoid clustering in a sparse dimensional space and to improve clustering performance ( Dalmaijer et al., 2021 ), we then projected the five SDQ subscales into a lower dimensional space using UMAP. This method is based on Riemannian geometry and algebraic topology and has been show to outperform Principle Component Analysis, t-distributed Stochastic Neighbor Embedding, or Multidimensional Scaling methods for (a) its ability to retain higher order, nonlinear interaction of variables inherent in the type of data we are studying; (b) its preservation of both local and global structure of the original data; (c) its greater flexibility with several distance metrics to choose from; (d) its computational efficiency (O[N] vs. O[Nlog(N)]); and (e) more reproducible results ( Yang et al., 2021 , p. 20).
Briefly, UMAP first constructs a topological representation of the original data with local manifold approximations, and then it optimizes the lower dimension embedding by minimizing the cross-entropy between the new lower dimensional space and the original higher dimensional space. We set UMAP parameters with a relatively high number of neighbors ( n _neighbors = 50) to decrease the probability of producing fine-grained cluster structure that is more prone to be a result of patterns of noise. We set a low minimum distance (min_dist = 0.01) to make denser clusters and cleaner separations. We also used a correlation distance metric. UMAP was implemented using the UMAP V.0.3.2 Python implementation ( McInnes et al., 2020 ).
We then used k -means clustering on the UMAP data in the lower dimensional space using the scikit-learn implementation in Python. Optimal number of clusters were chosen based on silhouette coefficients exceeding 0.5 and a steep increase in the Calinski–Harabasz index, both indicating a good separation between clusters scores ( Bathelt et al., 2021 ; Dalmaijer et al., 2021 ), and Jaccard similarity coefficients between clusters obtained after bootstrap resampling (1,000 iterations) and the original clusters ( Zumel & Mount, 2014 ) exceeding 0.85, indicating a good stability of clusters.
The non-elevated portion of the data was included as a comparison group for each age point. Comparisons between clusters for gender, cognition, mental health, and neurodivergence diagnosis were conducted with chi-square tests and one-way analysis of variance with Bonferroni post hoc t tests. All clusters (including the comparison non-elevated group) were compared to the whole sample at each age point.
Assessment of Transitions
Proportion z tests (Bonferroni-corrected for multiple comparisons) were used to compare the proportion of participants transitioning from one cluster at one time point to another cluster at the next time point with an equal-split transition (e.g., in the case of five clusters, an equal split would be 20% of the sample transitioning into each cluster). Only significant transitions above the equal-split proportion, that is, transitions that occurred more than we would expect by chance, will be reported here. To make sure transitions were not driven by higher Euclidean distance from the center of each identified cluster, we compared the distribution of silhouette scores for each participant who was part of a significant transition to the distribution of silhouette scores for the entire cluster (see Supplemental Figures 4 and 5 ). In other words, borderline cases could switch between clusters over time not because of any meaningful change in behavior but because of imprecision in the clustering solution. Comparing the silhouette scores in this way controls for this.
Identification of Longitudinal Risk Factors for Significant Transitions
Longitudinal and concurrent predictors of significant behavioral transitions between the ages of 5–11 years old and between the ages of 11–17 years old were identified looking at the feature importance gain after performing XGBoost on our train samples (80% of our data set) using the XGBoost Python package. For each significant transition, we compared the subgroup of children with a particularly elevated difficulty profile transitioning to another elevated difficulty profile cluster at the following time point (“transition of interest”) to the subgroup of children with the same elevated difficulty profile transitioning to the non-elevated difficulty scores group at the following time point (“control transition”). In essence, we were looking to identify factors that distinguish those who transition to a different profile of behavioral difficulty from those who transition to no apparent difficulties, given the same starting profile. XGBoost is an ensemble algorithm based on gradient-boosted trees. It integrates the predictions of “weak” classifiers to achieve a “strong” classifier (tree model) via a serial training process. It incorporates an L1 (Least Absolute Shrinkage and Selection Operator) regression regularization, which prevents the model from overfitting and is robust to multicollinearity of features.
We first removed all participants missing more than 30% of the features ( Uh et al., 2021 ; 282 and 178 participants were excluded for the “5–11” cohort and the “11–17” cohort, respectively), and after splitting our data set into training and testing sets (80%–20% as we prioritized low variance of our parameter estimates rather than our performance statistics), we imputed missing data using k -nearest neighbor imputation and scaled features using Min–Max method (3.35% and 3.30% of data were imputed for the 5–11 cohort and the 11–17 cohort, respectively). We then tuned the following hyperparameters optimizing both accuracy and f 1-score using a GridSearch procedure with 5-fold cross validation on our training sample: nrounds (the number of sequentially built trees in the model), η (the learning rate), max_depth (the maximum levels deep that each tree can be grown), min_child_weight (the minimum degree of impurity needed in a node before a split is attempted), γ (the minimum amount of splitting by which a node must improve the predictions), and reg_alpha (the L1 regularization term). We also set the scale_pos_weight hyperparameter with the inverse of the class distribution to tune the behavior of the algorithm for imbalanced classification problems. We then identified the optimal feature importance score threshold and number of features, which led to the best model performance accuracy and f 1-score for predicting the transition of interest in our test sample. Permutation importance was finally implemented on the test data by sequentially replacing each feature with its random shuffling and measuring at which extent permutation of a particular feature affected the accuracy of predictions. Only the features that survived both thresholding and permutation importance were considered genuine risk factors for a significant behavioral transition. We chose to report feature importance gain in this article, which corresponds to the improvement in accuracy brought by a feature to the branches it is on.
The data and code that support the findings of this study are available from the corresponding author upon reasonable request.
Transparency and Openness
We report all data exclusions, all data transformations, and all measures in the study, and we follow the Journal Article Reporting Standards ( Appelbaum et al., 2018 ). All analysis code is available at https://github.com/eliacello/MCS_BehaviouralTransitions_2021/tree/main . The code in this repository was used to clean, preprocess, cluster, and model data from the Millennium Cohort Study, known as “Child of the New Century,” and freely accessible through the UK Data Service ( University College London, UCL Social Research Institute, Centre for Longitudinal Studies, 2024 ). The Millennium Cohort Study data are not included in the repository. Data were analyzed using R, Version 4.0.0 ( R Core Team, 2020 ), and the packages EFA.dimensions Version 0.1.8.4 ( https://CRAN.R-project.org/package=EFA.dimensions ), ggplot Version 3.5.1 ( Wickham, 2016 ), polycor Version 0.8-1 ( https://CRAN.R-project.org/package=polycor ), and ez Version 4.4-0 ( https://CRAN.R-project.org/package=ez ). Data were also analysed using Python, Version 3.10, and the libraries numpy, pandas, matplotlib, seabird, sklearn (preprocessing, metrics, cluster), scipy, umap, and statistics. This study’s design and its analysis were not preregistered.
Figures, tables, references, and supplementary files are best inspected in the licensed PDF or repository copy linked above.