Fetching from the wire…
Research2026-09-07 · source-backed
Hiding continuous output probabilities does not close the door: a logit_bias parameter can be mathematically manipulated to evaluate exact probability thresholds with strictly one query per sample (arXiv 2609.05125). The authors build a provably consistent estimator of True Calibration Error for binary tasks on top of it. As a builder this is a working recipe for auditing calibration of a black-box model you only reach through an API.
Each link below shares sources, entities, or timing with this story.
Puro-2B trains from scratch on up to 1.4 trillion tokens in FP8 on consumer GPUs, approaching Qwen2.5-1.5B under the authors' protocol, against a stated $1.5M+ to train Llama-3.2-3B and $700K+ to reproduce SmolLM3-3B. The savings stack rather than coming from one trick: hardwa...
A July 2 evaluation tested semantic chunking against simple approaches on long structured academic theses using RAGAs, and the sophisticated method didn't win. Performance varied more with document formatting, preprocessing, and query type than with chunking strategy. The auth...
arXiv 2609.03376 targets organizations outsourcing vector indexes to untrusted clouds, where each query touches corpus-scale state, so a naive secure implementation costs minutes and about 90 GB of communication per query at million-document scale, and recent optimized systems...
arXiv 2609.04024 tests a black-box control that just duplicates the procedural instruction. Across seven instruction-tuned models, 300 medical multiple-choice questions, eight placement conditions and 16,800 scheduled generations, going from one copy to two raised the determin...
arXiv 2609.04172 trains OPD with a single query and finds it keeps improving for hundreds of steps across task domains and model families. Measuring state coverage, the fraction of full-data states a query set's rollouts reach, one query hits 71.5% and 16 semantically distinct...
Task-Conditioned Least-Privilege Learning trained Qwen3.5-4B on 1,500 terminal and MCP tasks to choose an authority level that fits the task, audited by deterministic verifiers across six risk dimensions. Over 2,896 episodes on 500 held-out tasks, excess-authority errors fell...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.