Research organisation
ResearchCore Ltd
Research team
Tori Olphin, Dr Tim Lowe, Fiona McDowell
Project start date
09/06/2025
Funding
£94,094
Primary dataset(s) used
Local held police and other linked administrative data
Status
Ongoing
Sectors
Policing, Youth Justice, Children’s Services
Why we funded this
Previous evidence has identified factors associated with an increased risk of children and young people’s (CYP) involvement in violence. Less is known about how these factors combine, and how administrative data can be used to predict risk and identify cohorts of CYP who would benefit most from support. This research will fill this gap, supporting better targeted service delivery, more effective testing of interventions and more efficient delivery of support.
What are the main research questions
The main research questions this project will address include:
- Which indicators (on their own or in combination) from locally held or other administrative data are the best predictors of children and young people’s involvement in crime and violence?
- Do risk and protective factors differ between children and young people from different ethnic backgrounds, of different ages, and of different genders?
- Is there an additional benefit from the use of machine learning to predict these outcomes, compared to rule-based models?
What the analysis involves
Datasets will incorporate police information, alongside local authority information to provide richer information about risk and protective factors.
Data will be matched between systems and then outcomes will be examined for individual factors, and for combinations of factors. Factors will be mapped against real world data to allow mapping of risk and protective factors, and identification of needs, and these will be examined for CYP from different ethnic backgrounds, of different ages, and of different genders to assess whether risk and protective factors are similarly impactful across different demographics.
Tools will then be built using a range of methods, including: single-factor predictiveness; rule-based models for simple cohort identification; and, machine learning techniques. The aim will be to examine whether more complex methods provide better identification of risk, or whether simpler tools can provide close to predictive accuracy.