BekchiAI Releases 2,057 Verifier-Checkable Agent Tasks Where Gold Answers Are Computed, Not Written
BekchiAI pairs a benchmark of 13 tool-using ReAct agents across 7 categories (arithmetic, structured/SQL, security detection, URL grounding, planning, orchestration, tool-policy) totalling 2,057 deterministic committed tasks with a web observability and control layer for live agents. Every gold answer is computed by running canonical SQL against a real database, computing an exact DAG schedule, or evaluating closed-form lambdas, and adversarial security samples are deliberately paired with imperfect signature scanners so a score reflects the model's own judgment rather than copying an oracle. It reports tool-call adherence, URL hallucination and source-match, and per-model token cost across Qwen3.7-Max, gemma-4-31B-it, gemma4:26b and gpt-oss-120b, with benchmark, evaluation tools and platform publicly released.
↳ Follow the thread