Fetching from the wire…
Skills2026-04-11 · source-backed
The 60% performance regression in cuBLAS dispatches the wrong kernel for all batched FP32 workloads on RTX GPUs. Profile your local inference with nsys or ncu to see if you're hitting the simt_sgemm_128x32_8x5 kernel path. If so, you're leaving 40-60% performance on the table.
Each link below shares sources, entities, or timing with this story.
Shared entity: MachineLearning / Same source domain / Earlier coverage / Downstream implication
Both cover MachineLearning; reported by the same outlet (reddit.com); earlier MachineLearning coverage from 2026-03-14.
Same source domain / Shared topic / Downstream implication
Reported by the same outlet (reddit.com); overlapping topics (gpus, inference, local); traces where this leads (what it means).
Shared entity: Check / Shared topic / What happened next
Both cover Check; overlapping topics (against, check); picks up the Check thread on 2026-08-07.
Both cover Check; overlapping topics (check, path); picks up the Check thread on 2026-08-03.
Shared entity: MachineLearning / Same source domain / What happened next
Both cover MachineLearning; reported by the same outlet (reddit.com); picks up the MachineLearning thread on 2026-07-29.
Shared entity: FP32 / Shared topic / What happened next
Both cover FP32; overlapping topics (against, performance); picks up the FP32 thread on 2026-07-26.
Shared entity: Check / Shared topic / What happened next
Both cover Check; overlapping topics (check, kernel); picks up the Check thread on 2026-04-30.
Shared entity: Check / Same source domain / What happened next
Both cover Check; reported by the same outlet (reddit.com); picks up the Check thread on 2026-04-19.