Fetching from the wire…
Research2026-09-05 · source-backed
arXiv 2609.03721 tested Gemini 3 Pro, DeepSeek R1, Kimi K2 and Qwen3 on 30 developer-written architectural design decisions from open-source projects, scored with ROUGE-L, BLEU, METEOR and BERTScore plus manual review of Gemini's outputs. Few-shot improved alignment (Gemini 0.828 to 0.847). Manual review found the generated ADDs were too long, implementation-focused, and missing the rationale, which is the only reason to write an ADD. A high similarity score on a useless output is the finding (arXiv).
Each link below shares sources, entities, or timing with this story.
The first systematic study of deceptive UI impact on LLM web agents, accepted at IEEE S&P 2026, tested against real e-commerce, streaming, and news dark patterns. Gemini 2.5 Pro: 65.78% susceptibility. Claude 3.7 Sonnet: 53.79%. GPT-4o: 51.26%. Guardrail models and prompt post...
Accuracy drops 30–50% well before you hit the documented context limit. Not at the limit. Before it. Cross-model testing across GPT-4.1, the Claude 4 family, Gemini 2.5, and Qwen3 quantified what everyone shipping long-context features has felt and couldn't measure (Glasp). Th...
Deng et al. built 120 real-case-grounded tasks across 20 business scenes in six financial domains, running four self-evolving scaffolds on a shared Qwen3.7-Max backbone against paired non-evolving controls. Letta posted the highest evolved score (91.65) and fewest compliance i...
LLM-Stats logged a cluster of coding-model releases on June 12 alone, continuing the month's heavy cadence. Chinese labs keep pushing fast, cheaper, code-specialized checkpoints into the same week as OpenAI's GPT-5.5 family and Google's Gemini 3.1 line. The practical effect: t...
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
The February rankings reshuffled: Windsurf #1 (Arena Mode for side-by-side model comparison), Antigravity (Google) #2, Cursor #3 (8 async subagents + Multi-Agent Judging), Kimi Code NEW at #4 — the first open-source tool in the top 5 with 100-agent swarm capability backed by K...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.