Dispatch
Ora benchmarks six agent harnesses against live customer sites and measures native success versus web-search fallback
Vercel published Ora's methodology on August 21, 2026. Ora sends Claude Code, ChatGPT, Gemini, Hermes, OpenClaw and Vercel's Eve onto real customer websites with tasks like signing up, integrating, and completing payment, then records where they fail. Per agent it tracks steps to completion, success rate with a specific split between native success on the site and fallback to web search, valid endpoints discovered and callable, plus cost and latency. Each harness gets an isolated runtime for side-by-side traces, run by a 16-person engineering team shipping hundreds of commits a day.
Source
↳ Follow the thread