Open APA research · Launch collection CC BY 4.0

Dynamic prediction of reoffending in individuals given community sentences: Development and validation of a novel risk monitoring assessment tool (oxMore).

Yukhnenko D, Blackwood N, Lichtenstein P, Fazel S.

Law and human behaviorAmerican Psychological Association2025-12-04DOI 10.1037/lhb0000641

Abstract

Objective This study aimed to develop and validate a dynamic risk assessment tool for individuals serving community sentences that accounts for the effects of acute adverse events and desistance from crime. Hypotheses Dynamic risk prediction models that incorporate updated data on mental health relapses, incidents of victimization, and desistance from crime will produce more accurate risk stratification for reoffending than models lacking dynamic measurement of risk factors. Method We analyzed a national cohort of 59,676 individuals given community sentences in Sweden, of whom 23,879 (45%) had prior psychiatric diagnoses and 18,546 (31%) had substance use disorder diagnoses. Model development tested prespecified criminal history, sociodemographic, and clinical risk factors. Employing landmarking methods for time-to-event data, we modeled the effects of new health care episodes during community supervision, changes in a supervised individual's circumstances, and the impact of crime desistance. We validated the model in a geographically distinct population. Results During follow up, 18,307 (31%) were reconvicted, 4,416 (7%) committed a violent offense, and 5,381 (9%) were hospitalized with a psychiatric diagnosis. The model demonstrated strong calibration and discrimination performance (c-index = 0.74 for violent reoffending, c-index = 0.69 for general reoffending). It also outperformed comparison models that did not incorporate dynamic data. The final model was translated into an online risk calculator (OxMore). Conclusions Implementation of dynamic models could lead to more accurate risk stratification for individuals under community supervision, including those with psychiatric and substance use disorders, potentially improving resource allocation, and linkage to interventions that reduce recidivism rates. (PsycInfo Database Record (c) 2026 APA, all rights reserved).

Attribution and reuse record

Authors
Yukhnenko D, Blackwood N, Lichtenstein P, Fazel S.
Original journal
Law and human behavior
Publisher
American Psychological Association
Publication date
2025-12-04
DOI
10.1037/lhb0000641
License
CC BY 4.0
Open repository
Europe PMC · PMC7618717
Collection
School leadership launch collection

Presented by the Journal for School Superintendents under the license identified in the article’s open full-text record. The original authors and publisher do not endorse this journal or its agent.

Open full text

Read the scholarly record

Method

The present study is a model development and validation study that uses a retrospective cohort design. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis checklist guidelines for model development and validation were followed (see Appendix A ; Collins et al., 2015 ). This study was approved by the Regional Ethics Committee at the Karolinska Institutet (2013/5:8; Stockholm, Sweden) and carried out according to the relevant guidelines and regulations. Written informed consent from participants was not required as the study was conducted on anonymized routinely collected register data and received ethics approval on this basis.

We linked the data from several longitudinal, nationwide Swedish registers: the National Crime Register, encompassing data on criminal offenses and convictions dating back to 1973; the National Patient Register, providing psychiatric diagnosis information for inpatient hospitalizations (since 1973) and outpatient care (since 2001); the Migration Register, detailing migration dates to and from Sweden; the Cause of Death Register, with dates and causes of death since 1958; the Multigenerational Register, providing insight into familial relationships for individuals residing in Sweden since 1933; and the Longitudinal Integration Database for Health Insurance and Labor Market studies, supplying annual estimates of income benefit receipt, marital and employment status, and education since 1990. The data linkage utilized a unique personal identifier allocated to all residents and immigrants in Sweden within national registers ( Ludvigsson et al., 2009 ).

Participants

We included Swedish residents aged 18 and above who had received any community sentence between January 1, 2007, and December 31, 2013. We excluded individuals born before 1958, as they would lack an uninterrupted criminal record in the National Crime Register, given that the age of criminal responsibility is 15, and criminal data were available from 1973 onward. Moreover, individuals who committed an offense before the beginning of the follow-up period but had not been sentenced for it by that time (referred to as pseudoreconviction) were omitted. Individuals whose cases had been appealed or dismissed were also excluded, as their information was not present in the sentencing register.

The community sentences comprised probation with community service, probation with contracted treatment, and conditional sentences mandating community service. These sanctions covered all types of community-based penalties in Sweden, delineated in Chapters 27 and 28 of the Swedish Penal Code ( Boijsen & Tallving, 2017 ), excluding postcustodial supervision and probation concurrent with imprisonment. The initiation of the follow-up period for each individual was determined by the date of receiving a community sentence. In cases where an individual received multiple community sentences during the study period, we randomly selected the index sentence and used its date as the starting point for follow up. This method ensured the inclusion of both individuals undergoing their initial community sentence and those with prior criminal records who had multiple community sentences. As a result, we ensured that the analysis cohort was more comprehensive and representative than the cohort consisting exclusively of individuals serving their first sentence.

Outcomes and Censoring

We developed separate models for two primary outcomes: the probability of violent reoffending and the probability of general reoffending, each within a 2-year window. General reoffending was defined as the commission of any offense following the index sentence, while violent reoffending included offenses such as homicide, assault, robbery, arson, sexual offenses, illegal threats, or intimidation. Offense dates were obtained from national registers and are recorded retrospectively once established by a court; if no offense date was available, the sentencing date was used as a proxy. Probabilities for both outcomes were estimated dynamically at sentencing and at 36 subsequent monthly landmarks during a 3-year follow-up period. A landmark refers to a predefined time point at which a new prediction is generated, using information available up to that time.

Censoring events were death from any cause, permanent emigration from Sweden, and, in the case of violent reoffending, imprisonment for a nonviolent crime. If an individual observation was censored before the end of the 2-year follow up, it was removed from the training and validation data.

Variable Specification

We categorized all predictor variables used in the analysis into four distinct groups ( Appendix B ). For the sake of coefficient interpretability, we did not consider interactions among covariates.

Group 1 included sex and previous criminal history covariates measured at the time of sentencing. All of the Group 1 covariates were included in all models by default. Group 2 consisted of covariates measured at baseline but subject to change during the follow-up period, reflecting an individual’s current status. This group included sociodemographic factors (that also included in Group 3), such as civil status, employment, receipt of income support, and housing stability, but measured dynamically and updated at each landmark using the most recent available data. In addition, Group 2 included current age and psychiatric diagnoses recorded prior to a given landmark. The variable “any psychiatric diagnosis/disorder” was defined using International Statistical Classification of Diseases , 10th revision diagnostic codes (F00–F99) recorded in the National Patient Register. This included individuals with or without a co-occurring substance use disorder (F10–F19). These covariates were included by default, as they capture time-varying aspects of risk that may influence reoffending trajectories during community supervision. Covariate values were updated every 30 days, starting from their initial values on the first day of the follow-up period, with these time points termed “landmarks” within the survival analysis framework using a landmarking approach (as described below). Since sociodemographic covariates were recorded annually (in November), we carried forward the most recent measurements until the next available record.

Group 3 consisted of covariates measured solely at baseline that were not updated during follow up but were thought, based on prior literature and theoretical relevance, to be associated with the outcomes. These included civil status (coded as single vs. not single, with the latter category including individuals who were married or in a registered partnership), employment status, receipt of income support, unstable housing (defined as more than three address changes within 1 year), and a recorded history of self-harm or suicide attempt. The inclusion of Group 3 covariates in the final models depended on the outcome of the variable selection process outlined below.

Group 4 comprised covariates exclusively measured during the follow-up period and likely associated with reoffending, such as triggers for violent crime ( Sariaslan et al., 2016 ) and psychiatric hospitalizations, serving as a proxy for acute and significant psychiatric symptomatology. We introduced the time-dependent impact of each trigger as three distinct binary variables, representing three hypothetical aspects of a trigger’s influence: the “acute effect,” “short-term effect,” and “residual effect.” The acute effect, indicating a risk surge, was assigned a value of 1 if a trigger event occurred within a week before a specific 30-day covariate update (landmark); otherwise, it was coded as 0. The short-term effect was assigned a value of 1 if a trigger event occurred within a month before the update; otherwise, it was coded as 0. The residual effect was assigned a value of 1 if a trigger event occurred at any time from the beginning of the follow up until a specific update; otherwise, it was coded as 0. Covariates for acute and short-term effects can be understood as modifiers of the residual trigger effect during their corresponding time windows.

Immutable personal characteristics such as sex and age are among the strongest predictors of recidivism and are included in most actuarial risk assessment instruments. In Sweden, the United Kingdom, the United States, and many other countries, their use is legally permissible and routinely incorporated into risk modeling. However, in some jurisdictions, their inclusion may raise legal or constitutional concerns, and their use should therefore be carefully considered in context, and may require adaptations (e.g., in training).

Handling of the Missing Data

The percentage of missing data for the included covariates varied from 0.1% to 3.2% at the time of sentencing, that is at baseline ( Appendix C ). We performed imputations for sociodemographic factors both at the baseline and during the follow up. In instances of missing sociodemographic factors during follow up, we utilized baseline values for the initial 3 months of the follow-up period, with 3 months representing the median time from the start of the follow up to the subsequent measurement. For education level, we extended the last recorded measurement until the next available measurement without any time restrictions. Missing values for other sociodemographic records were imputed using an expected-maximization algorithm implemented in the Amelia package for R ( Honaker et al., 2011 ), with the time of measurement serving as a cross-sectional time-series indicator. All measured covariates and outcome variables were utilized as predictors for missing data points ( Sterne et al., 2009 ).

Clinical covariates, including triggers, were recorded in the register only if the corresponding event had occurred. Consequently, we assumed complete information about clinical covariates in the data set and no imputation for missing values was deemed necessary.

Data Splitting

The data set encompassed the entire cohort, which was split into a derivation sample (approximately 80% of the data) used for developing and internally validating the predictive model, and an external validation sample (around 20% of the data). The derivation set was selected from the entire cohort based on the individual’s residential geographical location at the time of sentencing. Regions were primarily defined by the counties of Sweden, identified from the initial two digits of the Swedish Small Area Market Statistics code. Exceptions were made for the municipalities of Gothenburg and Malmö, which were distinct from their corresponding counties, and the Stockholm municipality, which was separated from its county and subdivided into northern and southern parts by associating each Swedish Small Area Market Statistics area with its historical province. These regions were categorized into four groups ( Appendix B ), serving as proxy indicators of urban/rural status.

This stratified geographic splitting strategy was designed to maximize heterogeneity between the derivation and validation samples while maintaining representativeness within each. By incorporating variation in population density and service infrastructure, it enabled a more rigorous evaluation of the model’s generalizability across different criminal justice settings. The external validation set was geographically distinct and selected randomly, with equal probability, choosing one region from each of the first three groups and proceeding sequentially through the fourth group. This methodology, previously utilized for other evidence-based tools, is recommended to avoid overfitting in the validation sample and to maintain cohort representativeness ( Fazel, Wolf, Larsson, et al., 2019 ).

Modeling Process

The predictive modeling employed Cox proportional hazards regression with sliding window landmarks ( van Houwelingen & Putter, 2011 ) adjusted for measured covariates. A “landmark” denotes a specific time point at which the values of covariates were updated; in our study, this occurred every 30 days of the follow up. At each landmark, we also recalculated the probability of the outcomes. Consequently, each landmark represented a time point for reevaluating the risk, and for each landmark, a distinct baseline hazard was estimated. In general, the landmark approach involves splitting the follow-up period into a series of overlapping time intervals and fitting separate risk models for each interval, using the information available at that time point. By then combining these models, the approach captures how risk changes over time in response to updated information.

In the analysis, we used 37 landmarks, starting from the commencement of the sentence (Landmark 0) and covering each month of the 3-year follow-up period (Landmarks 1–36). Individual data sets were created for each landmark, comprising only those individuals who were at risk at that specific landmark’s time, along with their corresponding covariate values. All 37 landmark data sets were subsequently consolidated into a single comprehensive landmark superset. The change in baseline hazard from one landmark to another landmark models the effect of desistance from crime since the start of the follow-up period.

To derive estimates of coefficients for the final prediction model, we combined results across multiple imputations using Rubin’s rule ( Barnard & Rubin, 1999 ). Rubin’s rule involves calculating the average of the parameter estimates across all imputed data sets to obtain a pooled point estimate. It then combines the within-imputation variance (reflecting sampling variability in each imputed data set) and the between-imputation variance (reflecting uncertainty due to missing data) to produce a total variance estimate. This approach yields valid standard errors, confidence intervals, and significance tests under the assumption that data are missing at random. By incorporating both sources of uncertainty, Rubin’s rule ensures that statistical inferences appropriately reflect the variability introduced by imputation.

Variable Selection

The variable selection process for the current model followed the general approach implemented during the development of the OxRec ( Fazel et al., 2016 ), FoVoX ( Wolf et al., 2018 ), and OxMIV ( Fazel, Wolf, Larsson, et al., 2019 ) prediction tools. All variables from Groups 1 and 2 were included in the final model on the basis of theory. To further select predictors from Groups 3 and 4, we employed an automatic backward elimination approach based on combined estimates from all imputed data sets ( Wood et al., 2008 ). We used a p value threshold of 0.157, which is equivalent to model selection using Akaike information criterion ( Sauerbrei, 1999 ), for excluding variables, following the convention to ensure model balance and prevent overfitting, as suggested in Heinze and Dunkler (2017) .

This approach ensures the face validity of the final model while concurrently permitting the inclusion of additional risk factors if they display an association with reoffending outcomes. The classification of variables into these four groups aims to produce as parsimonious a model as possible (i.e., easier to implement in practice), as long as it maintains acceptable predictive performance.

Several variables were reoperationalized after the first variable selection (see Appendices D and E ). The primary dynamic landmark model (DLM) for violent reoffending included: baseline factors (age, sex, and employment), being a victim of a violent assault, TBI or injuries from other causes, any psychiatric hospitalization, substance intoxication, any prior self-harm or suicide attempt. These factors were also included in the final model for general reoffending, although the triggers had different time components. DLM for general reoffending also included the receipt of income support at baseline as a predictor (see Appendix F for DLM formulae, coefficients, and baseline hazards).

Figures, tables, references, and supplementary files are best inspected in the licensed PDF or repository copy linked above.

Open Paper Agent