Fetching from the wire…
Public story · 2026-07-31 · high
It activates 12B of 276B parameters, ships with open weights on Hugging Face, and costs $1.20 per million output tokens.
Why now: Thinking Machines published Inkling-Small's numbers and pricing on July 30, one day before this coverage.
Thinking Machines released Inkling-Small on July 30, a coding model that activates just 12 billion of its 276 billion total parameters on every token, according to the company.
That's the headline number: 80.2% on SWE-Bench Verified from a model running a small fraction of its parameters per token. Inkling-Small also posted 31.6% on Humanity's Last Exam and 82.2% on IFBench, per the release.
The model runs a 1-million-token context window and works natively across text, image, and audio. Reasoning effort is adjustable per request, according to Thinking Machines.
Weights are open on Hugging Face, and fine-tuning runs through Thinking Machines' own Tinker service. Output tokens cost $1.20 per million, per the company. The release doesn't list a price for input tokens.
Open weights and Tinker fine-tuning mean teams can test Inkling-Small against their own coding workloads before committing budget. Whether the $1.20 output price holds up against real usage patterns is the open question the release doesn't answer.
Each link below shares sources, entities, or timing with this story.
Mira Murati's lab finally shipped a full LLM, and it's Apache 2.0. Inkling is 975B total parameters with 41B active in a MoE configuration, multimodal on input (text, image, audio) and text out, trained on 45 trillion tokens. The context number is the fun part: 1M tokens in th...
This is a supply-chain fact, and most people are still treating it as a geopolitics argument. Sequoia published "America's Open-Model Paradox" on July 24 with the number that reframes the whole conversation: Qwen's share of open-model fine-tunes went from 1% in January 2024 to...
Announced July 27 with Microsoft, IBM, Red Hat, Palantir, CrowdStrike, Cloudflare, Databricks, Hugging Face, LangChain, Nous Research, Reflection AI, Thinking Machines Lab, SpaceXAI and the Linux Foundation. Huang's framing is pointed: during the Hugging Face incident "closed...
1. Use claude agents --json to build session dashboards. Claude Code v2.1.145 outputs all live agent sessions as structured JSON with status, model, elapsed time, and parent relationships. Pipe it into a tmux status bar widget or session picker script for switching between bac...
DeepSeek released V4 Preview on April 24 with two open-weight variants: V4-Pro (1.6T total parameters, 49B activated via MoE) and V4-Flash (284B parameters, 13B activated). Both support 1M-token context windows. Both are Apache 2.0 licensed. Both are live right now on Hugging...
Zoph co-founded Thinking Machines Lab with Mira Murati and served as CTO, left in January 2026 with Luke Metz to return to OpenAI, was put on enterprise AI sales, and was later revealed to have been fired after five months (TechCrunch). Google to OpenAI to Thinking Machines to...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.