Fetching from the wire…
Public story · 2026-03-04 · source-backed
First cross-language benchmark for formally verified code generation: 40.3% success in Dafny, 24.7% Verus, 7.8% Lean. LLMs handle high-level verified code but collapse on systems-level constraints and manual proofs. (arXiv 2602.09464)
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source domain / Shared topic / What happened next
Both cover Dafny, Lean, Verus; reported by the same outlet (arxiv.org); overlapping topics (benchmark, lean, verified).
Terence Tao uses Lean / Shared entity: Lean / What happened next
Linked by a graph relationship (Terence Tao uses Lean); both cover Lean; picks up the Lean thread on 2026-08-02.
OpenAI uses Lean / Shared entity: LLMs / Same source domain / What happened next
Linked by a graph relationship (OpenAI uses Lean); both cover LLMs; reported by the same outlet (arxiv.org).
Shared entity: LLMs / Same source domain / Shared topic / What happened next / Tension
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (code, llms).
OpenAI uses Lean / Shared entity: Lean / What happened next / Tension
Linked by a graph relationship (OpenAI uses Lean); both cover Lean; picks up the Lean thread on 2026-08-03.
OpenAI uses Lean / Shared entity: LLMs / What happened next / Tension
Linked by a graph relationship (OpenAI uses Lean); both cover LLMs; picks up the LLMs thread on 2026-03-18.
Shared entity: LLMs / Same source domain / Shared topic / What happened next
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (benchmark, code).
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (code, constraint).