Reveree Scores Reverse-Engineering Agents Past Solve Rate and Finds Bigger, Newer Models Are Not Reliably Stronger
CTF solve rate says neither where an RE agent failed nor whether a success came from analyzing the binary or recalling a public writeup. Reveree scores the trajectory at three tiers, solve rate, milestone progress through an eight-stage RE schema, and a behavioral action profile, with comprehension stages judged by an outcome-blinded LLM validated against a human expert and everything else verified deterministically. Across nine frontier models and four prompting strategies on 88 picoCTF and NYU-CTF challenges, the base model dominates and prompting strategy is a secondary model-dependent effect, larger or newer or costlier models are not reliably stronger, and failures concentrate at comprehension where extra budget and reasoning effort rescue few.
↳ Follow the thread