An "Agentic Coding Index" aggregating seven coding benchmarks lands 142 upvotes and an immediate methodology objection
The index blends DeepSWE v1.1 (20%), Code Arena Elo (20%), Terminal-Bench v4.0 (15%), SWE-bench Pro (15%), Terminal-Bench v3.0 (13%), Terminal-Bench v2.1 (12%) and LiveCodeBench v6 (5%), then divides by sqrt(param count + an 8B regularization floor) raised to a 2.5354 super-linear exponent so tiny models cannot game the density score. The top comment asks the obvious question, how you get parameter counts for closed frontier models, and the answer is that API price is used as a proxy, which produces estimates like Claude Fable 5 at roughly 7.5T parameters from its $20/1M pricing. Useful as a shape, not as a number, since the denominator for every closed model is a guess derived from a business decision.
Source
↳ Follow the thread