67.9% of tool-retrieval queries have equivalent valid answers the benchmark marks wrong, making 30-47% of reported fine-tuning gains an artifact
Tool retrieval benchmarks annotate each query with one relevant tool combination, but repositories contain many functionally equivalent tools, so valid retrievals get scored as failures. ToolEX automatically discovers and annotates equivalent combinations; applied to the 7,360-query Tool-DE benchmark it found 67.9% of sub-queries admit alternatives, expanding the ground truth to an average of 5.3 valid combinations per query. Re-evaluating eight base retrievers and two fine-tuned variants on the expanded ToolEQ showed 30-47% of the reported fine-tuning gain was an evaluation artifact, and the same one-to-one problem reproduced on skill retrieval, which is directly relevant if you are tuning a skill or tool selector against a single-label set.
↳ Follow the thread