Fetching from the wire…
Agents2026-09-19 · source-backed
arXiv 2609.19391 freezes human-audited APIs and safety requirements as specifications, translates generated code into Dafny, repairs violations from verifier feedback, then compiles back to executable form. All 220 examples across 100 CUDA kernels, 100 terminal scripts and 20 robotic-arm tasks produced programs carrying non-trivial guarantees against the frozen specs. arXiv The paper's own caveat is the honest one: independent evaluations still found failures wherever auto-formalized semantics didn't capture intended behavior, which relocates the trust problem to specification quality rather than removing it.
Each link below shares sources, entities, or timing with this story.
After 20+ years maintaining Paint.NET, Rick Brewster concluded WINE's Direct2D would never be complete enough for what he needed, so the app now carries its own from-scratch reverse-engineered Direct2D implementation. He puts it at 180,000 lines against 700,000 for the rest of...
VAKRA (arXiv 2608.12282) benchmarks agents against 8,000+ executable APIs across 62 domains, verifying by re-executing predicted calls against live endpoints. Accuracy falls to 50-51% on compositional APIs and degrades over 50% as depth grows. Failures concentrate in entity di...
arXiv 2607.28187 black-box tested three foundation-model-based moderation services against seven model-agnostic transformations requiring no gradients or surrogate models. All three fall. Color inversion and grayscale conversion flip unsafe-to-safe while leaving content plainl...
Tsuchida et al. analyzed 622 GitHub users publicly signaling GenAI adoption, 179 repos carrying visible AI-assistance config against 179 matched traditional repos, plus 248 issues from the AI-assisted set. AI-assisted repos carry longer READMEs with more headers and code block...
First cross-language benchmark for formally verified code generation: 40.3% success in Dafny, 24.7% Verus, 7.8% Lean. LLMs handle high-level verified code but collapse on systems-level constraints and manual proofs. (arXiv 2602.09464)
For monolith-to-microservice translation, a graph pipeline built on tree-sitter ASTs, a Spanner property graph and a Gemini context cache dropped API hallucination from 56.4% to 16.2% and raised dependency resolution from 34.8% to 65.9%. Text-overlap metrics rated both at 91%,...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.