Skills
ModelScan is perfect on the model files it can judge and silent on half of them
A 170-artifact, 145-family benchmark (arXiv 2608.27424, 2026-08-27) scores three model-weight scanners on whether they return a decision at all, not just whether the decision is right. ModelAudit produced a definitive verdict on 100% of the 135 labeled families, Fickling on 81.5%, ModelScan on 49.6%, yet ModelScan hit 100% precision, recall and F1 on the subset it did judge. The operational lesson is to treat a scanner's N/A as unscanned rather than clean, and to check the coverage rate before trusting a published F1 on a supply-chain gate.
↳ Follow the thread