Skills
A 37.6ms local guardrail catches 95.9% of harmful prompts, against 255ms for Llama Guard 2
Reflex-Guard combines jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven binary classifiers, trained on 30,568 samples from five sources. It reports 95.9% recall at 37.6ms end to end, versus 255ms for Llama Guard 2 and 723ms for SafeDecoding, and 100% detection of GCG suffix attacks and Base64-encoded prompts at default settings. Running locally also removes the round trip to a cloud safety API, which matters when the prompt itself is sensitive.
↳ Follow the thread