Semantic Similarity Misses the Model Diversity That Actually Predicts Correlated Failure
Multi-model systems treat different models as independent components even when their failures stay strongly correlated, and existing diversity assessments use semantic similarity, which captures only differences in the meaning of observed outputs. The authors argue for generative-process diversity, differences between the processes capable of generating those outputs, measured via Normalised Compression Distance between raw model outputs residualised against a permutation control, drawing on Algorithmic Information Theory. Across 38 language models this measure identifies population structure semantic similarity misses and predicts chance-corrected correlated failure among model pairs across ten disjoint benchmark families, with a cross-benchmark partial rank association of -0.216 (95% interval -0.309 to -0.122) that is negative on all ten benchmarks, beyond semantic similarity and model-pair capability.
↳ Follow the thread