Fetching from the wire…
Agents2026-09-21 · source-backed
iApp Technology released OpenThai-SystemOne September 20, a 0.8B Apache-2.0 model with weights and the full training recipe, explicitly to open the architecture TypeSafe kept closed. On Bespoke Labs' 13-subset benchmark with identical subsets, splits, instructions and sampler, it scores 61.9 macro against 74.8 for Bespoke-Nimble-9B, 76.0 for Jev 1.13.0 and 45.4 for the raw Qwen3.5-0.8B base. Two independent open implementations in a week suggests the moat is the training data, not the architecture.
Each link below shares sources, entities, or timing with this story.
Jared Palmer published Kev September 17 on Qwen3.5, Apache-2.0 with training code and frozen eval suites. One request carries yes/no, multiple-choice and rating questions that share input text but can't read each other, and the API matches TypeSafe's System One so their Python...
Compaction is where long sessions go to die. The model summarizes what happened, the summary drops the exact string you needed three hours later, and you don't find out until the agent confidently references a file path that never existed. It's the biggest source of silent con...
@madiator's Bespoke Nimble uses contrastive data curation to lift the base model from 66% to 90% against Jev's 93% on the same task. Latent Space That's the most informative data point in the whole clone wave, because it suggests most of the accuracy is reachable with a public...
NandhaKishorM published it September 18 under Apache 2.0, a non-autoregressive decision engine evaluating typed questions (choice, score, noul) over text, email, tickets or JSON in a single forward pass, at 33ms for one question and 7.2ms per question batched, trained with RL...
The desktop app placed #5 on Product Hunt September 11 with 253 votes, positioned as "an open-source app for open-weight models" that runs multiple agent sessions and picks up work from Claude Code and Codex; desktop v0.0.25 arrived around September 10 and the Apache-2.0 repo...
SpeakoFlow Mini fine-tunes Qwen3.5-0.8B to apply only the corrections a speaker actually made and leave the rest alone. On the author's English-only benchmark it scored 70.7% against GPT-5.6 Luna's 65.0% under the same fixed short prompt with reasoning disabled, but the 95% in...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.