Research
RoSHAP: Distributional Framework Quantifies When SHAP Values Are Stable Enough to Trust
Introduces a distributional framework treating SHAP values as random variables rather than point estimates, plus a robust metric for measuring attribution stability across perturbations, resampling, and model retraining. Addresses a critical practitioner pain point: feature attribution methods often produce unstable rankings that shift with minor data or model changes, undermining trust in model explanations for high-stakes decisions. Provides actionable confidence bounds for when feature importance rankings are reliable.
Source
↳ Follow the thread