Domain-Adapted DarkBERT Beats Fine-Tuned LLMs at 89.33% F1 on Extracting Threat Intelligence From X
STINER releases a taxonomy and expert-annotated corpus of 2,100 real-world security alerts with eight entity types built around strategic pivots (Threat Actor, Sector, Location), benchmarking nine models across 12 configurations spanning general-purpose encoders, open-schema extraction, and zero-shot plus fine-tuned generative LLMs. The domain-adapted encoder DarkBERT reaches 89.33% strict F1, beating fine-tuned LLMs that also carry substantially higher inference latency — a useful counterexample to reaching for an LLM on every extraction task. Applied to a European H1 2025 threat analysis, the pipeline surfaced early signals of the SafePay ransomware campaign before its retrospective characterization in vendor reports.
↳ Follow the thread