Fetching from the wire…
Public story · 2026-09-10 · high
The chip's memory bus widens to 96-bit LPDDR5X, a change r/LocalLLaMA read as a bet on local AI decode speed on iPhone.
Why now: Apple announced the A20 Pro on September 9, and the bandwidth number is what r/LocalLLaMA focused on within a day of the reveal.
Apple's A20 Pro, detailed in the r/LocalLLaMA thread on the reveal, is the first 2nm smartphone chip. It widens the memory bus from 64-bit to 96-bit LPDDR5X, moving bandwidth to about 115.2 GB/s, a 50% increase over the A19 Pro. Neural Engine cores double too, to 32, up from 16.
The bus width is what got r/LocalLLaMA's attention, not the core count. Running a language model on a phone is a memory-bandwidth problem, not a compute problem. Decode speed depends on how fast the chip pulls model weights from memory for each token, more than on raw operation count. Apple could have kept the narrower bus and added GPU cores instead, for a cheaper spec bump. It didn't.
Widening a memory bus on a 2nm process is not a cheap change. More bus width means more die area, more power routing, more validation work, all for a spec that mainly matters if you're running large models locally rather than sending them to a server. Doubling the Neural Engine core count fits the same story: more dedicated silicon for running models on the device itself.
The announcement doesn't say which models Apple expects to run at that bandwidth, or whether third-party apps get the same access to the wider memory path that Apple's own on-device features get. Developers building local inference for iOS won't know the real ceiling until someone benchmarks a specific model on the A20 Pro and reports tokens per second.
Each link below shares sources, entities, or timing with this story.
You noticed the Mac mini shortage. Here's what caused it. The Information reported, via Cult of Mac, that OpenAI has purchased tens of thousands of M5 Pro and M6 Mac minis plus M5 Max and Ultra Mac Studios over the past several months, running reinforcement learning and comput...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
The developer beta pairs a local Neural Engine model for real-time Swift suggestions with a cloud routing layer that hands heavier analysis to Claude, Gemini, or OpenAI. The agent simulates whole apps, writes and runs tests, inspects visual changes via live previews, and drive...
OpenAI posted first benchmark results for Jalapeño, its Broadcom co-developed inference ASIC, claiming 1.5x to 1.9x more throughput per kilowatt and 1.7x to 3.6x lower end-to-end latency than Nvidia GB200 and GB300 rack systems, measured on the SemiAnalysis InferenceX suite. T...
Announced September 1, blocking third-party ads and trackers before they load, implemented in the browser because iOS doesn't support extensions the way desktop does (Mozilla). Runs on Apple's WebKit Content Blocker API against EasyList. It doesn't block first-party ads, searc...
In an August 4 filing, Apple said 11 additional former employees beyond Chang Liu and Tang Yew Tan may be involved, alleging meetings about unannounced-product information, screenshots of confidential documents taken before OpenAI interviews, and retention of Apple work device...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.