Voices
Simon Willison flags the Astra harness gap and a possible long-context breakthrough before he has even tested it
Willison's September 3 note is deliberately provisional, he has no access yet, but he isolates the two things worth watching: the 99.9%-vs-62.7% spread between the custom and default ARC-AGI-3 harnesses, and his read that OpenAI 'may have vanquished one of the ongoing challenges with long context processing' given 96.3% accuracy at 512K-1M tokens. He also notes Astra does not sweep, Fable 5.1 still leads on Artificial Analysis's Intelligence Index, 66 to 61, at identical $10/$50 per million pricing. No pelican yet.
↳ Follow the thread