COVER Wraps Any Video Grounder With Distribution-Free Coverage Guarantees — No Retraining, No White-Box Access
Video temporal grounding has an unacknowledged label problem: re-annotate the same query-video pair and independent annotators mark moments overlapping by less than half on a large fraction of samples, so ground truth is really a distribution over intervals — yet every grounder returns one interval with no reliability statement. COVER is a post-hoc model-agnostic wrapper that calibrates a temporal nonconformity score quantile on held-out labels and widens predictions to contain the true moment with probability at least 1-alpha, with finite-sample distribution-free guarantees under exchangeability. It supplies two score families (boundary-widening for interval outputs, super-level-set for relevance signals) and held target coverage across three benchmarks and five grounders.
↳ Follow the thread