Fetching from the wire…
Research2026-09-10 · source-backed
arXiv 2609.09560 ran 30 professional developers and advanced students through equivalent tasks under traditional, AI-assisted and AI-led conversational conditions with repeated-measures ANOVA plus thematic analysis. AI-led cut completion 27% against traditional and 12% against AI-assisted, scored SUS 71.4 and NASA-TLX 55.5, and produced lower maintainability indices with more security vulnerabilities. The thematic analysis ties the security regression to perceived loss of control, meaning less transparency and less validation of what the model emitted. Small n, self-reported mechanism, but the direction matches everything else this week.
Each link below shares sources, entities, or timing with this story.
When the meter's running hot, the obvious move is a cheaper model that's actually good. Mistral shipped one. Devstral 2 (123B, modified MIT) scores 72.2% on SWE-bench Verified. Devstral Small 2 (24B, Apache 2.0) hits 68.0%. Both carry 256K context. Mistral claims 7x cost effic...
PortLLM claimed training-free, data-free transfer of LoRA patches onto updated base models, but only over short horizons and without theoretical grounding. This study runs 10 continual-pretraining steps on Mistral, Gemma, and Qwen and finds portability persists long-run, meani...
"Understanding the (In)Security of Vibe-Coded Applications" finds that LLM-driven app generation routinely outpaces security review, leaving common vulnerability classes in shipped code. (arXiv) If you build with agents daily, this is the empirical version of a thing you alrea...
The help-center article states plainly that standard Vibe users are not opted out of training by default while Enterprise customers are, and that the Vibe and API opt-out toggles are separate settings a user must configure independently. Nothing about the page is new; what's n...
The authors define goal-directed execution as four repeated behaviors: selecting goals, constructing task-relevant state, maintaining fidelity to higher-level objectives, verifying completion against the environment. Post-training Qwen3.5-122B-A10B on 363 long-horizon multi-to...
Mira Murati's lab finally shipped a full LLM, and it's Apache 2.0. Inkling is 975B total parameters with 41B active in a MoE configuration, multimodal on input (text, image, audio) and text out, trained on 45 trillion tokens. The context number is the fun part: 1M tokens in th...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.