Fetching from the wire…
Security2026-09-15 · source-backed
Overflip is an instability in compact DeBERTa-class classifiers trained at 512 tokens with bucketed relative positional encodings (arXiv 2609.15013). On a 100-prompt benchmark, five of nine widely used guardrails flip MAL to BEN as input lengthens, with flip rates from 8% to 92% and first flips between roughly 2.6k and 9.4k tokens. The malicious content stays intact; repetition just homogenizes token-level attention over repeated structure. Different mechanism from classic attention-dilution padding. Test your guardrail at 10k tokens, not at 500.
Each link below shares sources, entities, or timing with this story.
RAGAS-style evaluation checks correctness against a frozen snapshot, which means routine document updates and corrections can silently break production without moving a dashboard. This ASE 2026 paper defines 11 mutation operators perturbing at both the pre-chunk index level an...
Writer launched Palmyra X6 on August 13 with a number that should reset how you think about agent COGS: 52% lower average cost, 48% better speed, 10% better quality. The model is a post-training variation of Z.ai's open-source GLM-5.2. A US enterprise SaaS vendor built its fla...
New work shows that fine-tuning LLMs for security classification introduces token-level evasion vulnerabilities that standard in-distribution evaluation completely misses, because models inherit circuits while learning new semantics (arXiv). If you're shipping a fine-tuned mod...
Reflex-Guard combines jailbreak-aware preprocessing, compact sentence-transformer embeddings and seven binary classifiers trained on 30,568 samples, reporting 95.9% recall end-to-end against 255ms for Llama Guard 2 and 723ms for SafeDecoding, with 100% detection of GCG suffix...
Across five TTS methods and five benchmarks spanning medicine, law, finance, chat and creative writing: candidate generation kept improving with compute in every domain, but reward models correlated with actual quality at roughly ρ=0.12. Only candidate *fusion* consistently be...
SpecPath found 35 of 100 passing implementations broke when only the revision path changed, with aggregate accuracy looking identical across paths. Build your eval set from real multi-turn clarification threads with amendments and reversals. Path sensitivity is invisible to st...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.