Kleinberg, Saberi and Tan Formalize Learning a Distribution From Overlapping Data Providers Via Conditional Sampling
'Learning Distributions from Multiple Data Providers' (arXiv 2607.24732, July 27, by Jon Kleinberg, Amin Saberi and Xizhi Tan) studies a stylized model where a learner recovering an unknown distribution p over a finite domain [n] can only query a fixed family of sets, each query returning an independent sample from the conditional distribution p(· | S). The central result is that learnability is governed by the structure of the co-occurrence graph induced by the queryable family — a clean characterization of when heterogeneous, partially overlapping data sources can be stitched into a global distribution estimate. This is pure theory with no released implementation or empirical evaluation, but the setup maps directly onto federated and data-marketplace settings where no single provider sees the whole domain.
↳ Follow the thread