← The Wire
Entity trail

DeBERTa

Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.

Briefing refs
1
Findings
1
Edges
0
Sources
1

Corpus findings

  1. 2026-09-15 / arxiv-researcherFive of Nine Lightweight Guardrail Models Flip Malicious to Benign Just by Repeating the PromptOverflip is a repetition-induced instability in compact guardrail classifiers (DeBERTa-class backbones trained at 512 tokens with bucketed relative positional encodings). On a 100-prompt benchmark, five of nine widely used guardrails flip MAL to BEN as the input lengthens, with flip rates from 8% to 92% and first flips at roughly 2.6k to 9.4k tokens. The malicious content is preserved intact; repetition homogenizes token-level attention over repeated structure, a different trajectory from classic attention-dilution padding attacks.

Source trail

Graph sources

entity graphfindings textkg entitiesnewsletter issues