about
Properly Constructed ML Models Are Not Necessarily Biased: Neuroimaging Study (arxiv.org)
1 point by PaulHoule on May 27, 2022 | hide | past | pdf | 2 comments on HN

In plain words: Brain-scan models trained on data pooled from many hospitals were tested on people of different genders, ages, and races to diagnose Alzheimer's, schizophrenia, and autism. They stayed accurate and fair across every group, and clinical and genetic details sometimes helped accuracy.

Abstract · Bias in Machine Learning Models Can Be Significantly Mitigated by Careful Training: Evidence from Neuroimaging Studies

Despite the great promise that machine learning has offered in many fields of medicine, it has also raised concerns about potential biases and poor generalization across genders, age distributions, races and ethnicities, hospitals, and data acquisition equipment and protocols. In the current study, and in the context of three brain diseases, we provide evidence which suggests that when properly trained, machine learning models can generalize well across diverse conditions and do not necessarily suffer from bias. Specifically, by using multi-study magnetic resonance imaging consortia for diagnosing Alzheimer's disease, schizophrenia, and autism spectrum disorder, we find that well-trained models have a high area-under-the-curve (AUC) on subjects across different subgroups pertaining to attributes such as gender, age, racial groups, and different clinical studies and are unbiased under multiple fairness metrics such as demographic parity difference, equalized odds difference, equal opportunity difference etc. We find that models that incorporate multi-source data from demographic, clinical, genetic factors and cognitive scores are also unbiased. These models have better predictive AUC across subgroups than those trained only with imaging features but there are also situations when these additional features do not help.

Rongguang Wang, Pratik Chaudhari, Christos Davatzikos
arXiv:2205.13421 · cs.LG, eess.IV · submitted May 26, 2022 · updated Jan 30, 2023
abstract · pdf · html

add comment on HN

Just read that one-paragraph summary, but... yeah, duh? I thought it was common knowledge, at least within the ML field, that network biases come from skewed training data, not from an issue with the model's architecture?
It's pretty hard to make a training set which is truly representative of the data.

For instance if you are collecting images of what could be a cancerous organ there is going to be a different prior distribution of cancers if you (1) sampled the population at random, (2) sampled people who had symptoms, (3) sampled people that following one set of guidelines for mass screenings, (4) sampled people that following a different set of guidelines.

It's quite problematic if people train a model based on one of the above and really use it on a different one.

Since cancers are rare compared to non-cancers you might want to enrich the data with more cancers to get a really good representation of cancers, but then that throws the prior distribution out of whack. It would be great to have a way to tune the prior distribution separate from the conditional probabilities as in Bayesian modeling but most ML approaches don't really do that.