Research
Summarization Metrics — Including LLM-as-Judge — Fail Basic Perturbation Tests of Informational Content
This paper argues the missing axis in summarization evaluation is whether a summary satisfies a specific reader, and that a persona (role and expertise) is a more practical signal than a query since users rarely state everything relevant — a biomedical researcher and a family doctor reading the same vaccine literature need different summaries. Testing sensitivity of popular metrics to both informational and persona differences, the authors find many, including strong LLM-as-judge metrics, fail basic perturbation tests of informational content. An expert human evaluation of information satisfaction agrees poorly with both traditional and LLM-based metrics.
↳ Follow the thread