Benchmark Radar Ships a Daily-Updated Catalog of 1,283 AI Benchmarks With 12,916 Numeric Score Observations
arXiv 2609.11115 (submitted 10 Sep 2026) describes a living database and search engine for AI benchmarks covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety and domain evaluations. Daily discovery draws on 37 sources (13 direct connectors and 24 first-party research and engineering feeds) and the catalog holds 1,283 source records from 4 benchmark catalogs with 12,916 numeric observations on 790 records, retaining source identities and citations plus mentions in model cards and technical reports. The paper audits the full catalog and walks a prior-art search end to end, which makes it a practical lookup when you need to know whether a benchmark is already saturated before quoting a score.
↳ Follow the thread