Fetching from the wire…
Research2026-08-11 · source-backed
arXiv 2608.07222 argues existing scaling laws systematically under- and overestimate loss at both data-scarce and overtraining extremes because they treat capacity and data as independent. One coupling exponent fixes it, cutting mean absolute percentage error 1.5–3x across interpolation and extrapolation. Paired with a sparse grid strategy, it achieves full-grid extrapolation at roughly 10x less compute than uniform sweeps. Reliable performance prediction before committing training budget is worth more than the accuracy gain.
Each link below shares sources, entities, or timing with this story.
It scales the fixed hidden state of gated linear RNNs by orders of magnitude by replacing Gated DeltaNet's dense key-value outer product with sparse reads and writes to a large explicit memory (arXiv). Under isoFLOP and equal-parameter constraints, the bigger state markedly im...
This study investigates dropping speculative decoding's lossless guarantee without any training, quantifying speed-ups against controlled capability drift. Standard spec decoding exactly preserves the sampling distribution. Relaxing it buys latency at a small distributional co...
Chimera processes text, image and video tokens as one raster-ordered stream with no positional embeddings, combining Kimi Delta Attention for O(N) state tracking, interleaved Multi-head Latent Attention, modality-aware short convolutions, and sparse MoE. The real contribution...
A new paper proposes gradient-based hallucination detection using internal model signals instead of external fact-checking or sampling-based consistency. The notable bit: it works at inference without needing multiple sampled generations, so the overhead is low enough to actua...
RGA-Designer trains a reward model scoring both task correctness and structural compactness, then fine-tunes a graph generator against it to design communication topologies. arXiv For fan-out agent teams where inter-agent chatter dominates the bill, topology is a cost lever mo...
arXiv 2608.12990 from Dongfang Li, Baotian Hu, Min Zhang and colleagues replaces turn-level memory consolidation with semantic boundary detection, reporting 89.22% on LoCoMo and 92.20% on LongMemEval-S while cutting construction tokens 86.0% and 75.9% versus the A-Mem baseline...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.