Fetching from the wire…
Public story · 2026-07-21 · high
The catch: scores now hinge on prompt wording, so two teams could land on different answers.
Why now: The paper landed in the July 21 arXiv batch with enough detail on the method to stand on its own.
Adobe researchers built an image similarity metric that reads a text prompt, per a paper posted to arXiv by Sheng-Yu Wang, Yotam Nitzan and Aaron Hertzmann. That matters for anyone doing image eval or dedupe work at scale. Fixed metrics like LPIPS force one similarity judgment onto every job, whether you care about shape or color.
Their fix: a metric that takes a text prompt specifying the axis of comparison. Instead of asking how similar two images are, you ask how similar they are in color, or in shape. The answer comes back scoped to that one question.
Right now, teams using LPIPS to catch near-duplicate images or score generative output are stuck with whatever notion of similarity the metric's training baked in. There's no way to tell it to ignore color and just check silhouette.
The paper doesn't say how the metric scores against LPIPS on standard benchmarks. It also doesn't address whether prompt phrasing shifts scores enough to make results hard to reproduce across teams.
A related paper tackles a similar problem: cutting agent-written code slop by judging the trajectory instead of the diff. Both point to the same failure: a single scalar number can't capture what people actually mean by good.
Each link below shares sources, entities, or timing with this story.
Same source
Cite the same source (arXiv).
Same source domain / Semantically similar
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.66).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.64).
Same source domain
Reported by the same outlet (arxiv.org).
Reported by the same outlet (arxiv.org).
Reported by the same outlet (arxiv.org).
Reported by the same outlet (arxiv.org).
Reported by the same outlet (arxiv.org).