Fetching from the wire…
Security2026-08-20 · source-backed
Reflex-Guard combines jailbreak-aware preprocessing, compact sentence-transformer embeddings and seven binary classifiers trained on 30,568 samples, reporting 95.9% recall end-to-end against 255ms for Llama Guard 2 and 723ms for SafeDecoding, with 100% detection of GCG suffix attacks and Base64-encoded prompts at defaults. (arXiv 2608.17556) Running locally also means the sensitive prompt never makes a round trip to a cloud safety API, which is the argument I'd lead with.
Each link below shares sources, entities, or timing with this story.
Shared entity: Guard / Same source domain / Shared topic / Earlier coverage
Both cover Guard; reported by the same outlet (arxiv.org); overlapping topics (attack, detection).
Shared entity: Running / Same source domain / Shared topic / Earlier coverage
Both cover Running; reported by the same outlet (arxiv.org); overlapping topics (cloud, combin).
Shared entity: Llama Guard / Same source domain / Shared topic / Earlier coverage
Both cover Llama Guard; reported by the same outlet (arxiv.org); overlapping topics (attack, guardrail).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (against, attack, compact, detection); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (against, attack, catch); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (against, attack, detection); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (against, binary, classifier); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (against, attack, guardrail); pushes against this story (against).