Research
JudgeStealer Replicates a Black-Box LLM Judge Across Three Evaluation Protocols From Pointwise Queries Alone
arXiv 2608.26982 presents a query-efficient model extraction framework aimed specifically at LLM judges, exploiting strong cross-protocol agreement to acquire pointwise scores and convert them into pairwise and listwise supervision without spending extra victim queries. It selects pointwise inputs by semantic diversity, predictive uncertainty and potential judge biases, then applies score smoothing and multi-protocol review to preserve ordinal structure during surrogate adaptation. Against state-of-the-art LLM-as-a-judge and reward models it reaches 73.3% pointwise, 87.0% pairwise and 71.6% listwise accuracy, holding up across surrogate scales and against representative extraction defenses.
↳ Follow the thread