Reddit
A 35B-A3B Tool-Calling Benchmark Found Ornith 1.5 and Tiel-Coder Tied, and That Tiel Is an Imatrix Quant of Ornith, Not a Finetune
A r/LocalLLaMA user ran tool-eval-bench across Qwen3.6-35B-A3B and its variants on a cluster of 32GB V100s, finding Ornith 1.5 and Tiel-Coder tied at the top, well above Qwen3.6-27B and close to 3.8-27B, with KAT-Coder slightly ahead of base 35B-A3B and Ornith-1.5-Heretic disappointing. The top comment corrects the framing: Tiel is an Ornith imatrix quant with a different chat template, not a separate finetune, though its author argues the gap shows up on SWE-Bench Live. When asked, the author also ran Qwen-AgentWorld-35B-A3B and got 121.0 average across five runs, about ten points behind base.
Source
↳ Follow the thread