Fetching from the wire…
Markets2026-09-01 · source-backed
Computable GPU Index took #2 with 194 upvotes, publishing USD-per-GPU-hour for H100, H200, B200 and B300 from 28 providers every 15 minutes, with both collection code and calculation methodology on GitHub so anyone can reproduce a published value (Product Hunt). It uses an interquantile mean so no single provider can move the print. It lands in a field where Silicon Data, GPUAlpha and gpu.ai sell exactly that opacity as the product.
Each link below shares sources, entities, or timing with this story.
Wafer.ai ran K3 at TP8 on a single MI355X node: 952 tok/s aggregate, 118 tok/s single-stream decode, ~13k tok/s steady-state prefill after fixing missing PyTorch sampling functions and AITER MLA prefill kernels via SGLang and ROCm. Against a two-node TP16 B200 deployment that'...
A practitioner ran Kimi K3 on 8x B300 via Modal at $56.79/hour with vLLM, TP8 and native MXFP4: 27-minute cold boot for a 1.56 TB load, TTFT 0.92 to 1.02s, 92 tok/s steady decode, roughly $36 of GPU time per clean run and $1,363/day left warm. The cheaper path was worse. Unslo...
The file is 1.56TB. That's the first thing you notice about the moonshotai/Kimi-K3 Hugging Face repo that went live today. 2.8 trillion total parameters, 104B activated, 896 routed experts with 16 selected plus 2 shared per token, 93 layers split 69 Kimi Delta Attention and 24...
Moonshot exposes an Anthropic-compatible endpoint, so pointing Claude Code at K3 means setting the Anthropic base URL and supplying a Moonshot key. No new CLI, no config rewrite. Hosted at $3/$15 per Mtok, same tier as Claude Sonnet 4.6, and Artificial Analysis scores K3 at 57...
The repo appeared on trending with +135 stars and a repositioned pitch, pivoting from the general local-code-execution tool it launched as in 2023. It's now aimed directly at Claude Code and Codex but on the open-weight side. Single-source on the repositioning, so check the re...
Hy4 preview, released August 28: 770B total parameters, 49B active, over 1M token context, open-sourced and simultaneously on Tencent Cloud TokenHub and OpenRouter at $0.834 per million input tokens, $2.501 per million output, $0.042 per million cached. In Tencent's own evalua...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.