Fetching from the wire…
Models2026-09-04 · source-backed
375B-A23B, 36B-A4B, 32B, 7B, 3.7B and 0.9B, sharing one architecture, vocabulary, training methodology and eval infrastructure. For every model they publish intermediate checkpoints, training data or the data-construction recipe, training code, configs, fine-grained logs and eval results, including the agentic post-training stages, which no other open family has exposed. IFM claims SOTA in class for the 0.9B, 3.7B and 7B tiers, with the 0.9B above 48 on AIME 2026, and a new Mixture-of-Value-Attention mechanism in the 36B-A4B. IFM
Each link below shares sources, entities, or timing with this story.
The top trending HuggingFace paper (274 upvotes) introduces dots.tts, a 2B continuous autoregressive text-to-speech model hitting best average Seed-TTS-Eval (WER 0.94%/1.30% zh/en) with strong cloning and emotional range. CFG-aware MeanFlow distillation gives 85ms first-packet...
GitHub | 196B total / 11B active, Apache 2.0 StepFun's sparse MoE activates only 11B of 196B parameters per token, delivering 74.4% SWE-bench, 97.3% AIME 2025, and 100-300 tok/s throughput. Supports INT4 GGUF for local inference. Apache 2.0 licensed. One of the most capable fu...
Triple-stream retrieval (BM25 keyword, vector embeddings, knowledge-graph traversal) fused via Reciprocal Rank Fusion on the iii engine, with SQLite for state and an in-memory vector index, no external database. The economic claim: ~170K tokens/year (~$10) versus ~650K tokens...
Cohere launched North Mini Code on June 9 under Apache 2.0, its first developer-focused model. The shape is the pitch: 30B parameters, mixture-of-experts, only ~3B active, and it runs on a single H100. It scores 33.4 on the Artificial Analysis Coding Index, competes on SWE-Ben...
XHToken published Spark-X2.5-4B and 1.7B with no announcement. Not a fine-tune: hybrid attention, one full-attention layer per three sliding-window layers, 200+ languages, ~20T training tokens (Hugging Face). Claimed 4B scores include 65.1 BFCL-V4, 75.1 tau-squared-bench, 44.4...
memoket/memoket-kite, created Aug 12, is an Apache-2.0 agent memory engine that drops embeddings, vector DBs, and rerankers entirely for a topic-indexed single file. README reports 93.51% on LoCoMo and 85.60% on LongMemEval-S using gpt-4.1-mini across all stages, average reade...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.