Fetching from the wire…
OSS2026-09-21 · source-backed
tinybrains.dev sorts submitted models into weight classes starting at 8 KiB, playing a reimplementation of Ants from the 2011 Google AI Challenge, scored for trained networks instead of hand-written bots. Entrants upload two files, a model and an adapter. 85 points on Hacker News September 20. The clearest recent example of a benchmark treating parameter budget as the constraint rather than the free variable.
Each link below shares sources, entities, or timing with this story.
Practical uncertainty quantification has to judge a single generation, but current methods either sample repeatedly, read only output-token probabilities, or collapse internals to one hidden state. ActMap compresses hidden-state trajectories across every layer and generated to...
Announced September 17, built and evaluated against real Wispr Flow dictations from offices, commutes and meetings across varied microphones and background noise. On 10 hours of real dictation it posted the lowest word error rate against Google, OpenAI, AssemblyAI and Deepgram...
Version 1.15.21, September 9. Any crewAI run trusting the framework's budget arithmetic for that model was overshooting by 56%. This is the argument against letting a framework own your context accounting: the number is a constant in someone else's repo and nothing tells you w...
A DiffusionGemma fine-tune writing UIs in openui-lang, a purpose-built format cutting token usage up to 67% against JSON and streaming progressively. It scores 71.7% on the Generative UI Benchmark against the base model's 13%, and on an unseen component library returned 55 val...
Artificial Analysis has the newly released Alibaba model at the top of its cohort: 27B reasoning model, 256K context, image input, Apache 2.0 permitting commercial use. The caveat in the eval data matters more than the headline: it emitted 160M output tokens across the benchma...
63 points on Hacker News September 11, a TypeScript repo created September 2 and now at 115 stars (GitHub). The bill of materials is the argument: Workers for compute, D1 for the ticket store, R2 for attachments, Queues for async. Zendesk and Intercom charge per agent for the...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.