Fetching from the wire…
Public story · 2026-08-10 · high
Manifest AI's one-line swap claims over 10x faster training and 100x faster inference at 64k context, MIT Technology Review reports.
Why now: MIT Technology Review published the profile August 10, when only two of the five projects, power retention and diffusion text, had working code ready to install.
Two startups already offer drop-in swaps for the transformer architecture that powers most large language models, per MIT Technology Review. The outlet profiled five firms chasing transformer alternatives, and three of the five are still research bets without a migration path.
Manifest AI's power retention layer replaces flash_attention with one line of code. Shipped as PowerCoder and Brumby. The company claims more than 10x faster training and over 100x faster inference at 64k context, according to the profile.
Subquadratic's SubQ uses sparse attention. It claims to be the first architecture to match top mainstream LLMs on search and coding tasks, per the outlet.
Liquid AI mixes 20% transformer with 80% of its own liquid neural network design. The result has 34 million downloads and matches rivals four times its size, including a version that runs on a $50 Raspberry Pi.
Inception's diffusion-based Mercury 2 generates GPT-4-class output about 10x faster and, like power retention, ships a working migration path already. Pathway's state-space Dragon Hatchling solved more than 97% of 250,000-plus sudoku puzzles in testing, a set where competing models solved none.
Only two of the five, power retention and diffusion text, have a working migration path. SubQ, the liquid hybrid, and Dragon Hatchling don't have a stated route to drop-in deployment, per the profile. For teams evaluating this, the two with working code are the ones worth testing first.
Each link below shares sources, entities, or timing with this story.
GPT competes with Claude / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (GPT competes with Claude); both cover GPT, MIT Technology Review; reported by the same outlet (technologyreview.com).
GPT competes with Claude / Shared entities / Earlier coverage
Linked by a graph relationship (GPT competes with Claude); both cover GPT, LLMs; earlier GPT coverage from 2026-03-17.
Copilot uses GPT / Shared entity: GPT / Shared topic / Earlier coverage
Linked by a graph relationship (Copilot uses GPT); both cover GPT; overlapping topics (becom, coding).
GPT competes with Claude / Shared entity: GPT / Earlier coverage / Tension
Linked by a graph relationship (GPT competes with Claude); both cover GPT; earlier GPT coverage from 2026-08-09.
Linked by a graph relationship (GPT competes with Claude); both cover GPT; earlier GPT coverage from 2026-08-06.
Linked by a graph relationship (GPT competes with Claude); both cover GPT; earlier GPT coverage from 2026-03-20.
Copilot uses GPT / Shared entity: GPT / Earlier coverage
Linked by a graph relationship (Copilot uses GPT); both cover GPT; earlier GPT coverage from 2026-07-20.
Linked by a graph relationship (Copilot uses GPT); both cover GPT; earlier GPT coverage from 2026-05-13.