CaSKG Calibrates Skill-Graph Edges With Counterfactual Probes, Beating Graph-of-Skills on All 12 Model-Benchmark Pairs
arXiv 2608.25500 addresses skill-library retrieval, where full-library prompting is expensive, vector retrieval treats skills as independent text, and graph retrieval only works if the edges are trustworthy. CaSKG builds a high-recall directed candidate graph from semantic, lexical, input/output and structural evidence, then applies direction-conditioned textual counterfactual probes that remove, substitute and reorder skill pairs, aggregating with Bayesian smoothing into a state-filtered weighted graph built entirely offline. Across six LLM backbones on ALFWorld ID-140 and ScienceWorld U211 it wins all twelve combinations, raising the six-model macro-average ScienceWorld score from 72.62 to 80.50 and ALFWorld success from 80.01% to 86.79% while cutting mean environment steps.
↳ Follow the thread