Fetching from the wire…
01
02
03
04
05
06
07
08
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
DeepSeek V4 Pro was evaluated on CommerceAgentBench at 53/107 tasks.
Source findingQwen 3.8 Max was evaluated on CommerceAgentBench at 53/107 tasks.
Source findingOpus 4.8 was evaluated on CommerceAgentBench at 56/107 tasks.
Source findingClaude Opus 5 was evaluated on CommerceAgentBench, leading at 65/107 tasks.
Source findingAlibaba International released the CommerceAgentBench benchmark.
Source findingGPT-5.6 Sol was evaluated on CommerceAgentBench.
Source findingDeepSeek V4 Pro was evaluated on CommerceAgentBench at 53/107 tasks.
Source findingAlibaba International released the CommerceAgentBench benchmark.
Source finding