Fetching from the wire…
Research2026-09-13 · source-backed
arXiv 2609.11023 inserts a layer between retrieval and generation combining a call-graph-derived structural coverage score with a novelty score estimating how far a query sits outside pretraining, then triggers targeted follow-up retrieval or flags for human review below a calibrated threshold. The argument is that on private codebases even oracle retrieval doesn't stop errors, it moves them downstream into API misuse, and model-internal confidence is the wrong signal for a private-code query. Evaluated by injecting synthetic internal APIs into open-source Java repos.
Each link below shares sources, entities, or timing with this story.
Data that contradicts the vibe. That's rare enough to lead with. Dipongkor, Baral, Lam and Moran analyzed 4,882 pull requests from five coding agents in the AIDev dataset (532 Java, 4,350 Python), accepted to ICSME 2026. The findings, in order of how much they should change yo...
Bridge mined 381,661 Java API update instances covering 18,900 mappings across 2,557 libraries, plus 277,259 Python instances, validated at 91.6% precision and 88.7% recall for Java (arXiv 2608.30497). Evaluated on replacement recommendation, the best model reaches 37.1% for J...
After 20+ years maintaining Paint.NET, Rick Brewster concluded WINE's Direct2D would never be complete enough for what he needed, so the app now carries its own from-scratch reverse-engineered Direct2D implementation. He puts it at 180,000 lines against 700,000 for the rest of...
AROMA+ automates the manual work behind Reproducible Central by recovering a library's source repository and original release environment from its Maven artifact, reaching up to 99.8% field-by-field accuracy against the hand-maintained list and catching flaws in it including b...
VAKRA (arXiv 2608.12282) benchmarks agents against 8,000+ executable APIs across 62 domains, verifying by re-executing predicted calls against live endpoints. Accuracy falls to 50-51% on compositional APIs and degrades over 50% as depth grows. Failures concentrate in entity di...
arXiv 2607.28187 black-box tested three foundation-model-based moderation services against seven model-agnostic transformations requiring no gradients or surrogate models. All three fall. Color inversion and grayscale conversion flip unsafe-to-safe while leaving content plainl...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.