Voices
Willison's One-Line Verdict on Muse Spark 1.2: 'The Most Important Characteristic of Any Model These Days Is Long-Sequence Agentic Tool Calling'
Reviewing Meta's release on August 5, Willison brushed past the benchmark table to name what he thinks actually differentiates models now — sustained tool-calling over long horizons, not single-turn reasoning quality. He ran his standard pelican-on-a-bicycle SVG test and called 1.2 'a small but material improvement' over 1.1, and added the new pricing to llm-prices.com. It is a useful evaluation heuristic to steal: if you are picking a model for an agent harness, the leaderboard score matters less than how many tool calls deep it stays coherent.
↳ Follow the thread