Fetching from the wire…
Public story · 2026-08-03 · high
SpyRL turns a hidden-identity voting game into the reward signal, with gains even on verifiable reasoning tasks.
Why now: RLSVR held the top spot on HuggingFace's Daily Papers list as of August 3, which is the news itself.
RLSVR climbed to the top of HuggingFace's Daily Papers list, drawing 138 upvotes there.
That's a signal worth watching for anyone training models on tasks without a checkable answer, like writing or summarization. Reinforcement learning has had no built-in way to score good output there.
RLVR, reinforcement learning with verifiable rewards, works for math and code because there's a checkable answer. Writing and summarization don't have one. RLSVR's fix, called SpyRL, borrows the structure of a social deduction game instead. Agents complete the same task under asymmetric information, then vote to identify an outsider. That vote becomes the checkable signal, standing in for a reward that open-ended writing can't produce on its own.
The gains show up on text summarization and creative writing, the tasks the method targets. They also show up on verifiable reasoning tasks the model wasn't trained to game directly. The paper treats that second result as the more surprising half.
The paper's bet is that skill at unmasking a spy in a hidden-identity game doubles as skill at writing well. That's an assumption, not a demonstrated fact, and a 138-upvote rank on one preprint doesn't test it against real readers.
Each link below shares sources, entities, or timing with this story.
The failure they target is specific and under-discussed: a cached error page or a negative price returns in the *expected schema* and gets consumed as fact, unlike a timeout the agent can see. Outcome Monitors check results against contracts mined from task-disjoint traces or...
Skill-α (arXiv 2608.01678) reframes skill generation as RL over sequential edits, decomposing skill construction into individually evaluable changes. The novel signal is a rollback reward that scores each modification by comparing downstream task execution using the original s...
It treats harness improvement as offline learning: diagnose failure traces, generate structured patches that edit the harness as source code, then select updates by validation over mini-batches of failures (arXiv 2608.23041). Gains: 9.0 on GAIA2, 9.6 on SWE-Bench Pro, 10.0 on...
Here's the number that should sit in every "agents will replace engineers" thread: 15.2%. That's the best model. The mean across 15 frontier models is 4.3%. Tencent Hunyuan's Long-Horizon-Terminal-Bench put 15 frontier models against 46 long-horizon terminal tasks across nine...
If you're tuning an agent system, upgrade the model doing the tuning before you rewrite a single line of the harness. That's the finding from HarnessOpt-Bench (arXiv 2608.06301, Scale AI), which tests whether frontier models can improve an agent *system* rather than write code...
arXiv 2607.27205 drops the standard LLM-centric V→L→A pathway for a direct V+L→A mapping with independent vision and language encoders and lightweight interaction. 31.2 ms latency, code released. Top-upvoted paper on HuggingFace Daily Papers today at 109 upvotes, and the first...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.