Audio-Language Model Benchmark: Gemini-3.1-Pro-Preview Leads at 85.6% Category F1 but Only 56.7% Fine-Grained — and States Wrong Answers Confidently 92–100% of the Time
Eleven audio classification methods — four Gemini models plus Kimi-Audio-7B-Instruct, four fixed-vocabulary taggers (YAMNet, PANNs, Whisper-AT, SSLAM), CLAP, and BAT — were evaluated in four tiers rather than one leaderboard over 2,242 clips across 23 fine-grained classes. Gemini-3.1-Pro-Preview led at 85.6% category-level and 56.7% fine-grained F1; Kimi-Audio reached 67.5%/32.9% but failed to answer 1.6% of samples, while SSLAM and CLAP matched or beat the best closed-set model at category level without seeing the candidate list. Analysis of 8,968 Gemini chain-of-thought responses found response length does not predict accuracy and the apparent 'holistic beats detailed' effect is a difficulty confound.
↳ Follow the thread