Fetching from the wire…
Research2026-08-02 · source-backed
Choukhmane, Lin and Akuzawa with Stanford's de Silva tested GPT-5.2, GPT-5.6 and Gemini 3 Flash. Advice produced sizable savings buffers for people over 30 but degraded on life-change adjustments and active rebalancing. Performance jumped substantially with structured prompts carrying explicit assumptions, because "regular people are not writing their prompts the way a finance professor is." That capability-versus-elicitation gap shows up in every domain, and it's the strongest argument for prompt scaffolding in consumer products.
Each link below shares sources, entities, or timing with this story.
Gemini built by Google / Shared entities / Earlier coverage
Linked by a graph relationship (Gemini built by Google); both cover Flash, Gemini, GPT, LLM; earlier Flash coverage from 2026-05-16.
Simon Willison released LLM / Shared entities / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover Flash, Gemini, GPT, LLM; earlier Flash coverage from 2026-07-31.
Simon Willison released LLM / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover Flash, Gemini, GPT; earlier Flash coverage from 2026-07-23.
Gemini competes with Claude / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Gemini competes with Claude); both cover Gemini, GPT, LLM; earlier Gemini coverage from 2026-03-20.
Simon Willison released LLM / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover Flash, Gemini, GPT; earlier Flash coverage from 2026-03-17.
Simon Willison released LLM / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover Flash, GPT; overlapping topics (gpt-5, prompt).
LLM uses OpenAI / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (LLM uses OpenAI); both cover GPT, LLM; overlapping topics (active, gpt-5).
Stanford benchmarked against DeepSeek / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Stanford benchmarked against DeepSeek); both cover Flash, GPT, LLM; earlier Flash coverage from 2026-06-15.