Machine learning to build a risk prediction model for lung cancer in never smokers
Lung cancer is a leading cause of cancer-related deaths, and early detection remains the key to improving survival. While lung cancer rates are declining among smokers, they remain stable in people who have never smoked — a group that currently has no access to screening programs. Existing prediction models for lung cancer in never smokers perform poorly and have not been validated in large populations.
This project aims to develop accurate risk prediction models for never smokers using the CARTaGENE cohort. Two approaches will be compared: traditional models based on known risk factors identified by experts, and machine‑learning models that can detect complex patterns in large datasets. Several advanced algorithms — including decision trees, neural networks, and ensemble methods — will be trained and internally validated to identify the best-performing model.
By improving risk prediction in never smokers, this research could help identify individuals who might benefit from future screening programs and guide public health strategies to reduce lung cancer mortality in this growing population.
- Cancer
- Environment and lifestyle
- Public health and epidemiology
- Biochemical and hematological data
- Genetic data
- Linked data
- Physical and cognitive measures
- Questionnaire data