Stack layer / Contrast
DFM Mimir v1: 1B-Parameter Hierarchical Reasoning Model Trained Entirely on Permissible Data, Competing With Qwen 3.5 4B
arXiv 2608.13517
Threat pattern / Contrast
Vertical Federated Learning Backdoor Results Collapse Under Realistic Constraints; BVBench Released to Reset the Field
arXiv 2608.12962
Stack layer / Update thread
Filtering multi-agent messages by answer correctness throws away value: over 4 in 10 outcome-changing wrong messages help
arXiv 2608.14375
Stack layer
'Compliance Theatre': A 120B LLM Judge Loses 47 Accuracy Points to Keyword Stuffing on Regulatory Evaluation
arXiv 2608.14329
Contrast / Update thread
E2-Explainer turns multi-agent communication topology from black-box optimization into a causal attribution problem — then prunes the graph
arXiv
Stack layer
'Beyond Final Scores' Instruments What Agents Actually Do Across 36 Long-Horizon R&D Tasks — Verdict: Engineering Optimizers, Not Researchers
arXiv (via HuggingFace Daily Papers)
Stack layer
A small RL-trained model at each node of a recovery graph diagnoses drift in autonomous LLM agents on AppWorld
arXiv
Stack layer
Thirteen LLMs tested on one-shot coordination: frontier models beat Nash, but the advantage vanishes at four or more agents
arXiv