Off-the-Shelf LLMs Translate Unstructured Requirements to Linear Temporal Logic Without Fine-Tuning, Across 450 Generated Formulas
Newcomb and Ochoa (2608.06287, submitted 2026-08-06) evaluate six modern LLMs on translating unstructured natural-language requirements into LTL formulas using only few-shot prompting — no task-specific fine-tuning — on a heterogeneous benchmark of 15 structurally varied requirements, collecting five independent generations per requirement-model pair for 450 candidate formulas total. They score with manual semantic evaluation plus pass@k for k in {1,3,5} and self-consistency measures, and conclude general-purpose models now reach practically significant performance on the task. The framing is deliberately modest — viable front-end assistants for semi-automated formalization workflows in safety- and mission-critical verification, not autonomous spec generation.
↳ Follow the thread