Fetching from the wire…
Public story · 2026-07-25 · high
The authors rank target emergence, goals shaped mid-interaction, above two other reasons humans stay in AI workflows.
Why now: The paper surfaced in coverage dated July 25, while most AI evals still assume a fixed target going in.
Fourati, Schütze, Hüllermeier and Gurevych argue humans stay in AI loops for reasons that don't reduce to model capability, per their preprint on arXiv. That matters for anyone building an eval, because a fixed-target benchmark can't score a goal that only forms mid-interaction. The authors say that covers most of the interesting domains in AI.
The paper names three grounds for that persistence: complementarity, normative and developmental value, and target emergence.
Target emergence gets the most weight in their ranking. It covers cases where the goal isn't set in advance and forms through the interaction itself.
Under target emergence, the human isn't cleanup crew for the model's draft. Participation is constitutive of the output, not a means of improving execution, per the preprint.
I don't sit down with Claude Code in my personal projects with a finished spec. The real target usually shows up three or four prompts in. After I've seen what the first attempt gets wrong.
The fix the authors point toward is an eval that lets the target move instead of locking it before the run starts. Nobody's built that yet, as far as I can tell.
Each link below shares sources, entities, or timing with this story.
Sampled softmax cuts the O(nK) memory of full-vocabulary classification to O(nk), but for fixed budget B = n·k it's been unclear whether to buy batch or negatives (arXiv 2608.11061). Analyzing convergence under standard smoothness and variance assumptions, the fastest converge...
Standard evaluation of frozen-embedding style classification uses random splits where works by the same artist appear on both sides (arXiv 2608.14435). Under an artist-disjoint protocol on 320 paintings across four twentieth-century movements, 5-NN accuracy falls ten points, a...
Learned KV eviction has a soft-to-hard mismatch: training uses differentiable gates that attenuate contributions, but inference only saves memory when entries are physically removed (arXiv 2608.23296). A controlled 2x2x2 over attention type, learned gating and positional encod...
arXiv 2608.06370 evaluated models emitting code that calls tools against JSON-schema tool calling on BFCL v4. PTC matched or exceeded the baseline in 11 of 14 models, with the GPT-5.6 family up 10.6%, and held stable under parallel execution in 13 of 14. Under context degradat...
arXiv 2608.03609 formalizes agentic systems over relational data as Stateful Tool-Enabled Agentic Deployments and proves verification against First-Order CTL specs is undecidable. Under a finite-domain restriction it becomes PSPACE-complete, but only if renaming opaque identif...
Studdiford and Lupyan tested human participants and 25 LLMs on common-sense causal reasoning and found shared, predictable error patterns triggered by irrelevant prompt details (arXiv). They localized the attention heads driving it. The uncomfortable implication: the gap betwe...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.