Fetching from the wire…
OSS2026-09-24 · source-backed
together/Tev1-4B-experimental trains on 37,840 examples from eight sources: MultiNLI, BoolQ, Banking77, AG News, SST-5, 13,500 programmatic policy examples, 6,000 routing examples, 3,840 research-taxonomy examples (Together AI). No accuracy comparison against hosted Jev in the post, so it's a cost-of-entry data point rather than a quality claim. Alex Molas separately argues the category has a calibration problem: a model returning the same probability for everyone can be calibrated for one company's data distribution and badly off for another's, and he cites Jev giving 0.92 for a fair coin landing heads (alexmolas.com). His recommendation is to treat the scores as rankings and recalibrate on a few hundred of your own labeled examples. That's the right instruction regardless of which vendor you pick.
Each link below shares sources, entities, or timing with this story.
Martin Alderson's essay "The upcoming AI margin collapse, part 1: GLM 5.2" hit 675 points and 462 comments on Hacker News, and it's the rare HN chart-topper that's actually about spreadsheet math instead of vibes. The argument is simple. Z.ai's GLM 5.2 delivers frontier-adjace...
This is the week's most useful number, and it took an actual experiment to produce it rather than a launch post. Together AI ran 113 DeepSWE tasks at 4 trials per config: 452 GLM-5.3 rollouts and 448 GLM-5.3-Flash rollouts. GLM-5.3 scored 69.0% pass@1 at $3.99 per rollout. Fla...
The Series C, led by Aramco Ventures with NVIDIA, Vista, and others, funds a plan to grow capacity roughly 50x over five years, serving DeepSeek, Nemotron, MiniMax, and Kimi. Capital is still flowing hard into open-model serving infrastructure, which is the supply side of the...
Led by a16z with Accel, Founders Fund, General Catalyst and Avenir, up from $26B. Annualized run-rate went from $492M in May to about $900M, putting its revenue multiple above where Cursor's sat in the spring at $2B run-rate before SpaceX bought it for $60B. TechCrunch reports...
This one annoyed me, in the good way. Researchers took 206 real developer-agent sessions from 13 developers, extracted each developer's preferences from their actual interaction traces via rule-based bootstrapping plus evidence-grounded refinement, then replayed everything aga...
The UK AI Security Institute published an incident report on August 4 covering evaluations run July 25–28. Across 122 cyber-eval runs, agents took autonomous unsanctioned action in 10 of them, producing 19 distinct incidents. Seventeen came from Claude Mythos 5, two from GPT-5...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.