Fetching from the wire…
Public story · 2026-09-06 · high
The same module built to personalize also catches 16.0-22.3% more failures than rubrics written by humans or models alone.
Why now: The paper published its results as agent builders debate whether personalization needs a massive user base to pay off, and TAHI's numbers argue it doesn't.
Researchers adapted an agent to 30 people across 600 tasks and raised solo success 4.5-20.9% within tens of tasks, per TAHI. That's the difference between an agent that stays generic and one that adapts fast enough to serve a single new user well.
TAHI folds a user's past sessions into both the agent's working context and its weights. A rubric module distills what a person wants from a task, then scores new attempts against that rubric. The 600 tasks spanned writing and visual creation, work where 'good' depends on who's asking.
The same rubric module doubles as a standalone annotation tool. Its rubrics catch 16.0-22.3% more failures than rubrics written by models alone or by humans alone, with no personalization step required.
The paper doesn't say how the gains hold up past 600 tasks or past the 30 people tested. It's unclear whether improvement keeps climbing as a user's history grows into the thousands.
Each link below shares sources, entities, or timing with this story.
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
First benchmark of off-the-shelf LLMs against expert-derived ground truth built on INCOSE criteria, ten models across two families and five generations each, one hundred independent runs, two requirement sets, five temperatures. The error profile is asymmetric, and performance...
ArcticSwarm separates evidence gathering from evidence integration: subagents publish to a shared board, but gated isolation lets selected search tasks keep their own prior so parallel agents stop converging on an early candidate before alternatives are tested. On full BrowseC...
The authors model strategic bidding as a repeated game with imperfect public monitoring, then run multi-agent RL over it, and build a criteria set for judging collusion that goes beyond comparing profit against Nash equilibria. Agents sustained supra-competitive outcomes match...
arXiv 2608.26733 presents an execution-only attack that reconstructs a hosted agent skill without ever asking the victim to reveal it, submitting crafted but ordinary tasks whose results discriminate between candidate hidden behaviors. At the weakest access level, final respon...
On a verifiable protein-function characterization task routed across tools, model choice swamped federation topology, RL-versus-LLM harness, and prompt expertise: Opus at roughly 92 to 94%, o4-mini at 40 to 50%. Federation across institutional boundaries cost almost nothing (a...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.