Fetching from the wire…
Research2026-09-16 · source-backed
Chain-of-Self-Questioning makes answer commitment conditional on an explicit assessment of what the question requires, with no training. On 817 TruthfulQA multiple-choice items, Grounded-CoSQ at tau=0.90 cut mean unconditional wrong-commitment from 13.1% under chain-of-thought to 8.9%, a 32.1% relative reduction, while raising answered accuracy from 86.9% to 89.7% and still answering 87.6% of questions. It held for all eleven models at every evaluated threshold.
Each link below shares sources, entities, or timing with this story.
Farid Zakaria's Self-Executing Linux Format uses binfmt_misc to hand the file to an interpreter that maps rows from a segments table and jumps to the entry point, with the program reading its own file via argv[0]. Symbols, relocations and application data all live in tables in...
He set the 4-byte SQLite application ID at offset 68 to "SELF", decomposed an ELF binary's components into rows across a custom schema, and registered a binfmt_misc handler that hands the file to a self-exec interpreter which queries the tables and runs the program. One file,...
FACE-Eval varies where a preference cue is delivered, user message or tool return, across 5,100 samples and 15 open-weight models from 4B to 1.60T parameters (arXiv 2608.29464). Every single model showed lower verbalized commitment for tool-return cues, and unverbalized adopti...
arXiv 2608.04735 points out that monitorability evals overwhelmingly use *explicit* influence, where the prompt tells the model to hide a side task, and monitors catch 60-94% across seven frontier extended-thinking models. Swap in subtle contextual bias and detection drops 41-...
DataSpace benchmarks data agents on 410 cross-language tasks over 7,439 artifacts totaling 15.01GB across CSV, JSON, SQLite, Markdown, PDF and video, validated by 11 domain experts. Six frontier multimodal models across five frameworks: best accuracy only 66.34%, and harness c...
BAAI's AREX (24 authors, 124 upvotes on HF Daily Papers) alternates between gathering evidence and drafting provisional answers, then audits those answers constraint-by-constraint. The distinguishing mechanism is a learned autonomous context-update tool that compresses growing...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.