Variable selection and estimation using expectile regression in the presence of measurement errors and missing data

Status Ongoing
Start date 2023
Location Quebec
Researcher Barry, Amadou
Associated institution Institut national de la recherche scientifique
Summary

This project aims to improve how researchers analyze large datasets when some information is missing or when certain measurements are not very reliable — two very common issues in health research. Traditional statistical methods do not handle these situations well, especially when the data are numerous and complex.
The team is therefore developing a new approach that looks for information where traditional methods rarely focus: in the “extreme” parts of a distribution, where people at higher risk for certain diseases are often found.
This method could improve the accuracy of polygenic risk scores, a tool used to estimate the likelihood that a person will develop a complex disease. It may also better represent populations that are often under‑included in genetic studies, such as some non‑European groups.
Part of the method has now been completed, and its application to CARTaGENE data is underway. Analyses are progressing, and the first scientific publications are being prepared.

Themes
  • Genetics and genomics
  • Methodology and biostatistics
  • Public health and epidemiology
Data types
  • Biochemical and hematological data
  • Genetic data
  • Physical and cognitive measures
  • Questionnaire data