Fetching from the wire…
Public story · 2026-09-09 · high
Pathway's Baby Dragon Hatchling skips chain-of-thought tokens entirely, updating memory during inference instead, at under a tenth of a cent per task.
Why now: AWS published the SageMaker HyperPod training writeup on the architecture on September 9.
Pathway's Baby Dragon Hatchling, a 150-million-parameter model, scored 29.2% pass@2 on ARC-AGI-1 at roughly $0.0007 per task, according to AWS's SageMaker HyperPod writeup on the project. That's a fraction of the parameter count and per-task cost most reasoning models run at.
The architecture, called BDH-CQ, skips the usual move of writing out chain-of-thought tokens. Instead it iterates in a recurrent latent state, with about 5% of neurons active at any point, and updates its internal memory during inference rather than through a separate fine-tuning pass. If that holds up, it's a different way of getting a model to reason step by step: state that evolves as it runs, not a transcript of tokens it has to generate and re-read.
Why it matters: ARC-AGI-1 was built specifically to resist the kind of pattern memorization that lets big models look smart on benchmarks without generalizing. A small model doing well on it, if the number holds, says something about efficient reasoning that parameter count alone doesn't capture.
The catch is that this is AWS's own blog post about a partner's model trained on AWS infrastructure. It's a vendor-published benchmark, not a third-party eval, and the post doesn't say whether anyone outside Pathway and AWS has reproduced the score. I'd treat 29.2% as the number to go verify, not the number to build a roadmap around. If independent labs run BDH-CQ against ARC-AGI-1 and land anywhere close to that figure, the cost-per-task math changes by orders of magnitude for anyone running reasoning workloads at scale.
Each link below shares sources, entities, or timing with this story.
A spec is a press release until someone who didn't write it implements it. GitHub made Agent Plugins 1.0 generally available on August 12 across VS Code, Copilot CLI, the Copilot SDK, and the Copilot app on all plans. The spec, published August 6, was co-authored by AWS, Anysp...
The full family, Sol, Terra, and Luna, is generally available on Amazon Bedrock with IAM and VPC controls (LLM Boss, AWS). Sol targets coding, biology, and cybersecurity agentic work. Terra runs everyday tasks at about half GPT-5.5's cost, and Luna optimizes for speed. The thr...
poetiq.ai — Detailed technical breakdown of how a $40K-hardware startup achieved 54% on ARC-AGI-2 (vs. Google's 45% at nearly 3x the cost). The key innovation: "learned test-time reasoning" — an iterative refinement meta-system where solutions are generated, receive structured...
First-party pricing, counts toward AWS commitments, Codex via CLI and IDE plugins for VS Code, JetBrains, and Xcode, across commercial and GovCloud (AWS). This removes the procurement and compliance wall for AWS shops that couldn't touch OpenAI under existing contracts. Distri...
NVIDIA's Blackwell successor is in production ahead of schedule. The NVL72 rack (72 GPUs) delivers 3.6 exaFLOPS for inference, with 288GB HBM4 per GPU. NVIDIA claims 10x lower cost-per-token versus Blackwell. The Rubin CPX variant — purpose-built for million-token inference —...
Built on a post-transformer BDH architecture that reasons recurrently in latent space, it hit 29.5% pass@2 on public ARC-AGI-1 at a computed cost roughly 11x cheaper per task than GPT-5.6 Luna Low, even after OpenAI's 80% price cut on July 30 (Pathway). Be precise about what t...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.