Fetching from the wire…
Security2026-09-19 · source-backed
An r/LocalLLaMA host rented their rig, observed a renter running vulnerability scans and attempting to post malware to a Colombian betting site over residential internet, and after Clore declined to cancel or block, mounted the container filesystem offline and recovered the scan logs, malware payload, reverse-proxy request-smuggling tactics and the AI agent reports generated along the way. r/LocalLLaMA One-sided account. The liability question for anyone monetizing an idle GPU at home stands regardless.
Each link below shares sources, entities, or timing with this story.
Cloudflare added Moonshot AI's Kimi K2.5 to Workers AI on March 19, making it the first frontier-scale open-source model available on edge compute with a full 256K context window, multi-turn tool calling, vision inputs, and structured outputs (Cloudflare Blog). Cloudflare repo...
Ollama cut v0.34.0-rc1 on September 5 at 23:49 UTC, and the headline item changes the shape of the local-versus-hosted decision rather than the performance of either side: Ollama-hosted open models can be selected directly inside ChatGPT Desktop, with setup driven from the Oll...
Up from $4,299.99 in June against a $1,999 launch MSRP, with Korean listings at $5,112 (r/LocalLLaMA). Memory is now over 80% of a GPU's bill of materials, with 16GB of GDDR7 climbing from about $65-80 per card in mid-2025 to over $200 by year end. Local inference economics ch...
AMD argues agents require so much CPU-heavy orchestration alongside GPU inference that the ratio is moving from 1:8 toward 1:1. If AMD is right, that means entirely new racks of CPU servers in every AI data center. Major implications for AMD's EPYC roadmap and Strix Halo APUs,...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.