Local Open-Weight SSH Honeypots: Prompt Design Beats Fine-Tuning, and the Two Fight Each Other
LLM-based SSH honeypots lean on closed cloud models for shell realism, which brings unstable versioning, provider-side changes, attacker-driven cost and eventual decommissioning. arXiv 2608.18686 (2026-08-19) fine-tunes and evaluates eight models, the original shelLM GPT-3.5 plus seven open-weight local models each against its base, using 34 automated unit tests for shell emulation accuracy in single-session and fresh-session settings. Prompt design has a large effect while fine-tuning depends entirely on dataset coverage: the original 112-conversation dataset does not improve aggregate pass rate, an expanded dataset built from real honeypot logs clearly does, and strong rule-based prompting can conflict with supervised adaptation because both address overlapping shell-behavior constraints.
↳ Follow the thread