Nature: AI Disease-Prediction Models Were Trained on Dubious Data
Nature·medium signal
A Nature investigation reveals that several widely-cited AI disease-prediction models were trained on datasets with significant quality issues, including mislabeled samples, demographic biases, and data leakage between training and test sets. The findings call into question the reliability of published accuracy metrics for clinical AI tools. For builders deploying health-adjacent AI, this is a reminder that model performance claims are only as good as the training data provenance.