Fetching from the wire…
Infra2026-09-14 · source-backed
Schemes that hide the model by returning noisy output-permuted responses backed by shuffle-model differential privacy fail inside the correctness regime they need. d+1 admissible queries to a d-input linear layer exactly recovers a permutation-invariant layer summary, which is enough for perfect model distinguishability. Demonstrated by recovering every linear layer of a ResNet-20 from TFHE transcripts with zero error in 5,712 queries.
Each link below shares sources, entities, or timing with this story.
arXiv 2609.09553 shows cipher-based covert-communication jailbreaks no longer need fine-tuning on an encrypted corpus. In-context learning is enough, and alignment is significantly weakened or bypassed once the exchange runs through the learned encoding. Demonstrated against m...
arXiv 2608.05604 names the mismatch precisely: current systems retrieve skills as packages but compress them as prose, which destroys the execution contract. SkillZip does contract-preserving compression over section-level graphs, rewriting recurring valid motifs into reversib...
In Tacet, an empirical analysis declares what it generated and what it expects to find, and is refused any claim it can't afford or properly test. A sample selected by reading outcomes permanently sets a purity bit and is recorded as having examined everything it read, so it c...
arXiv 2607.23438 argues autonomy debates conflate what an agent *can* do with what it *should be permitted* to do. AAL is authorized autonomy given risk, oversight, and accountability; ACL is inherent technical ability. The ladder runs reactive execution → decision support → s...
Across four benchmarks and 36 judge-examinee pairs, a model's own task accuracy predicts its judging accuracy at r ≥ 0.90, but capability doesn't buy fairness: more capable examinees get more lenient judgments from every judge at r ≥ 0.83. Calibrated weighted majority voting e...
First benchmark of off-the-shelf LLMs against expert-derived ground truth built on INCOSE criteria, ten models across two families and five generations each, one hundred independent runs, two requirement sets, five temperatures. The error profile is asymmetric, and performance...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.