An 8B Local Llama Beat Wazuh and OpenSearch on Auth Anomaly Detection, 89.3% Accuracy Against 52.0% and 49.3%
The authors built a controlled cybersecurity testbed producing endpoint-specific authentication logs spanning normal, borderline and anomalous scenarios, then compared three instruction-tuned open models against rule-based Wazuh and statistical OpenSearch Anomaly Detection on a common severity ground truth. Meta Llama 3.1 8B Instruct reached 89.3% accuracy, 88.2% recall, 91.8% F1 and an 11.8% false negative rate, versus Wazuh at 52.0% accuracy / 68.6% FNR and OpenSearch at 49.3% / 74.5%. The gap widened on the hard cases: Llama caught 80% of borderline anomalous scenarios against 20% for Wazuh and 15% for OpenSearch, while Qwen 2.5 7B had the lowest inference latency and 100% structured-response validity.
↳ Follow the thread