Fetching from the wire…
OSS2026-09-14 · source-backed
The project wires a CUDA application through ZLUDA to a cuBLAS/cuSPARSE/cuFFT shim to rocBLAS/hipBLASLt/rocSPARSE/HIP, using ZLUDA v6-preview.69, AMD HIP SDK 6.4 and LibTorch 2.3.0 against CUDA 11.8, validated by training a 2.2M-parameter PPO network. Only the RX 9060 XT (gfx1200) is verified, cuDNN isn't available in the stable Windows HIP SDK, and NCCL, TensorRT and custom CUDA extensions can fail. 110 stars and nine commits, so early, but it's the most concrete published CUDA-on-AMD-under-Windows recipe I've seen.
Each link below shares sources, entities, or timing with this story.
Reuters, via Tech Startups, reports capital released against deployment milestones with Anthropic deploying up to two gigawatts of Instinct MI450 starting 2027. Same structure as Nvidia/OpenAI: compute vendor capital flowing to the lab that commits to buy the silicon. A two-gi...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
+7,546 stars this week for an "Application Development Environment" running each concurrent agent in its own worktree. Works with any terminal CLI agent — Claude Code, Codex, OpenCode, Pi, Cursor, Copilot, Grok, 30+ others — across macOS, Windows, Linux, iOS via App Store/Test...
DeepSeek-V4-Flash-0731 landed July 31 under MIT with a DSpark speculative-decoding module attached. Terminal Bench 2.1: 82.7. Toolathlon-Verified: 70.3. DSBench-FullStack: 68.7. DeepSWE: 54.4. NL2Repo: 54.2. The model card claims it beats DeepSeek-V4-Pro (Preview) "despite its...
Announced September 4 and gated on developer-class hardware: 64GB+ unified memory and 250+ GB/s bandwidth, enough to run 30B-parameter models without paying per token. It preinstalls languages, runtimes and source control, pins Terminal and VS Code, turns on file extensions, h...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.