Fetching from the wire…
Public story · 2026-09-02 · high
XHToken's new 4B and 1.7B models claim benchmark scores beyond their size, and most local tools still can't load them.
Why now: The weights are already public, but the llama.cpp pull request is still open, so most local runtimes can't load the model yet.
XHToken published Spark-X2.5-4B and a 1.7B sibling with no blog post or launch thread, just model weights on Hugging Face under an Apache 2.0 license.
A fully open 4B model with a native 1M-token context window and no licensing restrictions changes what small local models can do. XHToken says the 4B version scores 90.7 on AIME 2026, one of four benchmark claims on the model card.
Neither model is a fine-tune of an existing base. XHToken built a hybrid attention architecture instead, one full-attention layer for every three sliding-window layers. It has native 1M-token context and training across more than 200 languages on about 20 trillion tokens.
XHToken's own benchmarks add three more figures for the 4B model: 65.1 on BFCL-V4, 75.1 on tau-squared-bench, and 44.4 on SWE-Bench Pro. All four scores are self-reported by XHToken.
The models don't run on stock llama.cpp yet. A pull request adding support, PR #27868, is still open. Until it merges, using Spark-X2.5 requires XHToken's custom fork or the pre-built GGUFs it ships alongside the weights.
Each link below shares sources, entities, or timing with this story.
DeepSeek released V4 Preview on April 24 with two open-weight variants: V4-Pro (1.6T total parameters, 49B activated via MoE) and V4-Flash (284B parameters, 13B activated). Both support 1M-token context windows. Both are Apache 2.0 licensed. Both are live right now on Hugging...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Published August 25, it's IBM's first family of dense decoder-only reasoning models, with the 30B flagship claiming state-of-the-art resolve rates on SWE-Bench Pro and Terminal-Bench. The 8B and 30B went through an agentic training curriculum on real sandboxes for software eng...
Released August 4 under Apache 2.0, reframing moderation as policy-adaptive question answering: write your rule in plain language, get a calibrated safety score from a single token, no retraining, one interface for text and images. Mistral claims it matches open guard models u...
Announced July 27 with Microsoft, IBM, Red Hat, Palantir, CrowdStrike, Cloudflare, Databricks, Hugging Face, LangChain, Nous Research, Reflection AI, Thinking Machines Lab, SpaceXAI and the Linux Foundation. Huang's framing is pointed: during the Hugging Face incident "closed...
1. Use claude agents --json to build session dashboards. Claude Code v2.1.145 outputs all live agent sessions as structured JSON with status, model, elapsed time, and parent relationships. Pipe it into a tmux status bar widget or session picker script for switching between bac...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.