Skills
Using past execution trajectories to pick the eval subset beats fitting pass/fail matrices
Efficient SWE-benchmark evaluation normally picks a representative subset by fitting historical pass/fail response matrices or static task semantics, discarding how agents actually solved anything. PTA-IRT feeds process-level evidence from historical trajectories, explored context, attempted edits and solving paths, into item response theory as privileged information for both subset selection and ability estimation. Under low calibration budgets it beat prior IRT baselines on score and ranking recovery across four SWE benchmarks, which matters if you are running a full benchmark on every agent change and paying for it.
↳ Follow the thread