Fetching from the wire…
Models2026-09-14 · source-backed
The report targets sandbox-to-production mismatch with a dual-track pipeline, triple-system consensus evaluation, and an error-correction module that salvages every failed trajectory into supervision rather than discarding it. 87.4 on MobileGUI-VBench, 5.1 points above the best closed-source model, and the best open-source AndroidWorld result reported.
Each link below shares sources, entities, or timing with this story.
CADWorld is a 200-task FreeCAD benchmark across 11 mechanical-CAD workflow categories, with agents operating through screenshots and GUI actions and success determined by executable checks over the saved native project. Seven current agents, best result 17.5%. The failure prof...
The Apache-2.0 repo was last modified September 16 with 1,048 downloads against paper arXiv:2609.11977. It's a Qwen3.6-35B-A3B post-train, and the pitch is cost-per-episode rather than peak capability: co-work agents spend most steps on state tracking, coordination, recovery a...
AgentHijack deployed patches on author-controlled GitHub Pages and a locally hosted CSDN clone against five GUI-agent backends. Across 600 online cases: 84.5% success at the VLM output stage, 47.0% at action parsing, 20.3% end-to-end with verifiable environmental consequences....
The PlayCoder benchmark is the cold shower the vibe coding movement needed. Researchers tested 10 state-of-the-art code LLMs on generating GUI applications across six categories. The models achieved high compilation rates. The code built and ran. But when they measured whether...
The 397B MoE scores 86.1 on Terminal-Bench 2.1, 86 on SWE-Bench Verified, 56 on DeepSWE, 92.8 on GPQA Diamond and 44.6 on HLE, which the team frames as comparable to Claude Opus 4.8. The method is a closed self-improvement loop where the model proposes its own tasks and scaffo...
RAGAS-style evaluation checks correctness against a frozen snapshot, which means routine document updates and corrections can silently break production without moving a dashboard. This ASE 2026 paper defines 11 mutation operators perturbing at both the pre-chunk index level an...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.