Elliot Arledge: Opus 5.5 at xhigh writes a Kimi-Linear decode megakernel at 35.5x PyTorch, ahead of GPT-6 Astra's 24.8x
The Neuron / X (Elliot Arledge)·low signal
KernelBench-Mega maintainer Elliot Arledge reported that Opus 5.5 at xhigh produced a single-kernel Kimi-Linear W4A16 batch-1 decode megakernel on an RTX PRO 6000 Blackwell that ran 35.5x an optimized PyTorch baseline. GPT-6 Astra reached 24.8x in the same unlimited-budget run. The benchmark rejects multi-kernel Triton pipelines, so this measures real fused-kernel authoring. The only source is Arledge's post as relayed by The Neuron.