Fetching from the wire…
Public story · 2026-07-19 · high
Red Hat is serving the reasoning model on DGX B200 through vLLM, putting it in a production inference stack instead of a lab demo.
Why now: The scores surfaced in coverage dated July 19, as Red Hat moves Inkling into a production vLLM deployment on DGX B200.
Thinking Machines' Inkling model posted 79.5% on ARC-AGI-1 and 36.5% on ARC-AGI-2, per Latent Space's July 19 coverage.
ARC-AGI-2 is the harder of the two benchmarks, the one where every lab's score has stayed low. Latent Space calls 36.5% meaningful movement in a spot where movement has been rare. That's the number enterprise buyers evaluating reasoning models should watch, not the ARC-AGI-1 score.
Red Hat is deploying Inkling on DGX B200 hardware, served through vLLM, the open-source inference server. That puts a reasoning-focused model into a standard enterprise inference stack. It's the same kind of setup enterprise teams already run production inference on, not a research demo.
Yes, but 36.5% still means Inkling misses roughly two of every three ARC-AGI-2 problems.
The coverage doesn't say how this score compares to Thinking Machines' earlier releases. It also doesn't say whether Red Hat's DGX B200 deployment is a pilot or a standing commitment.
Each link below shares sources, entities, or timing with this story.
Mira Murati's lab finally shipped a full LLM, and it's Apache 2.0. Inkling is 975B total parameters with 41B active in a MoE configuration, multimodal on input (text, image, audio) and text out, trained on 45 trillion tokens. The context number is the fun part: 1M tokens in th...
Zoph co-founded Thinking Machines Lab with Mira Murati and served as CTO, left in January 2026 with Luke Metz to return to OpenAI, was put on enterprise AI sales, and was later revealed to have been fired after five months (TechCrunch). Google to OpenAI to Thinking Machines to...
Published July 30: 276B total parameters with only 12B active, 1M-token context, open weights on Hugging Face. 31.6% on Humanity's Last Exam, 80.2% on SWE-Bench Verified, 82.2% on IFBench. Natively multimodal across text, image and audio, variable reasoning effort, $1.20 per 1...
This is a supply-chain fact, and most people are still treating it as a geopolitics argument. Sequoia published "America's Open-Model Paradox" on July 24 with the number that reframes the whole conversation: Qwen's share of open-model fine-tunes went from 1% in January 2024 to...
Founding signatories include AWS, Anthropic, Google, OpenAI, NVIDIA, Microsoft and GitHub, IBM, Red Hat, Cisco, JPMorganChase, Citi, the Rust Foundation, Zscaler, and Sonatype, with OpenSSF, CNCF, and OpenInfra participating. The open letter drew 455 points on HN. (Akrites / L...
The project (pure C, Apache-2.0, 71 stars) lazy-fetches only the bytes an inference touches and caches them locally, sending 4 KB activations to peers holding the relevant experts rather than transferring expert weights. Local and remote paths share identical code to guarantee...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.