Skip to content

Secondary Data Analysis

Predicting children and young people’s risk of involvement in violence

This project will identify factors associated with children and young people’s involvement in violence, and explore how new methods could better identify at-risk cohorts for more targeted support.

Research organisation

ResearchCore Ltd

Research team

Tori Olphin, Dr Tim Lowe, Fiona McDowell

Project start date

09/06/2025

Funding

£94,094

Primary dataset(s) used

Local held police and other linked administrative data

Status

Ongoing

Sectors

Policing, Youth Justice, Children’s Services

Why we funded this

Previous evidence has identified factors associated with an increased risk of children and young people’s (CYP) involvement in violence. Less is known about how these factors combine, and how administrative data can be used to predict risk and identify cohorts of CYP who would benefit most from support. This research will fill this gap, supporting better targeted service delivery, more effective testing of interventions and more efficient delivery of support.

What are the main research questions

The main research questions this project will address include: 

  • Which indicators (on their own or in combination) from locally held or other administrative data are the best predictors of children and young people’s involvement in crime and violence? 
  • Do risk and protective factors differ between children and young people from different ethnic backgrounds, of different ages, and of different genders? 
  • Is there an additional benefit from the use of machine learning to predict these outcomes, compared to rule-based models?

What the analysis involves 

Datasets will incorporate police information, alongside local authority information to provide richer information about risk and protective factors. 

Data will be matched between systems and then outcomes will be examined for individual factors, and for combinations of factors. Factors will be mapped against real world data to allow mapping of risk and protective factors, and identification of needs, and these will be examined for CYP from different ethnic backgrounds, of different ages, and of different genders to assess whether risk and protective factors are similarly impactful across different demographics. 

Tools will then be built using a range of methods, including: single-factor predictiveness; rule-based models for simple cohort identification; and, machine learning techniques. The aim will be to examine whether more complex methods provide better identification of risk, or whether simpler tools can provide close to predictive accuracy.