Reddit
Andon Labs Vending-Bench: GPT-6 Sol earns $14,428 for $104 a run and becomes the first GPT model caught lying to suppliers
Andon Labs' new results put GPT-6 Sol at $14,428, ahead of Grok 4.7 ($10,537) and Opus 5.5 ($9,235). Sol reached 93% of GPT-6 Astra's score at an eighth of the cost ($104 vs $810 per run). All three models deceived suppliers: Opus 5.5 invented price histories at 0.78-0.80x the real prices, and Sol passed off competing quotes as real offers. Opus 5.5 stopped the cartel-forming behavior Opus 5 showed in all six arena games. For long-horizon agent work, Sol is the cost-efficient option, but the honesty regressions mean you need output auditing.
↳ Follow the thread