Fetching from the wire…
Public story · 2026-09-09 · high
The Pro model beats Qwen3.8-Max on a computer-use benchmark by nearly 10 points, and all three sizes ship under Apache-2.0.
Why now: As of Sept. 9, Apache-2.0 weights for a 397B computer-use model are posted on Hugging Face.
Nex put out three computer-use models under an Apache-2.0 license: Mini at 35B parameters, Pro at 397B, and a text-only Max at 1.6T. All three, including the Mini weights, are posted on Nex's model card.
Apache-2.0 lets a team self-host, fine-tune, or resell any of the three sizes without asking Nex first. That matters most for Pro and Mini, the two that can actually run outside a hosted API.
Mini is a 35B mixture-of-experts model with about 3B parameters active per token, 262K context, and 256 routed experts plus a shared one. It weighs in at roughly 70GB in BF16. Two H100s run it, which puts a computer-use model within reach of one workstation instead of a cluster.
Pro, at 397B, carries the numbers worth arguing over. It scores 56.4 on OSWorld-2, ahead of Qwen3.8-Max's 46.7. On OSWorld-G it posts 87.4; on OSWorld-Verified, 82.2; on Vision2Web, 68.2. All four are Nex's own reported scores, not an independent run.
Max drops multimodal input for a 1.6T text-only design and scores 50.2 on AutomationBench v1.0.6, 0.1 point behind Claude Opus 5. Close enough to be noise, not a clear win either way.
What's missing: training data, inference cost per task, latency. A benchmark score on screen navigation doesn't say how the model handles a real desktop full of dialog boxes and pop-ups.
Each link below shares sources, entities, or timing with this story.
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
DeepSeek released V4 Preview on April 24 with two open-weight variants: V4-Pro (1.6T total parameters, 49B activated via MoE) and V4-Flash (284B parameters, 13B activated). Both support 1M-token context windows. Both are Apache 2.0 licensed. Both are live right now on Hugging...
Alibaba published a fine-grained MoE with 2.4T total / 95B active, 512 experts, and a 92-layer hybrid full/linear attention backbone. vLLM shipped day-0 support verified on NVIDIA and AMD with ready 4-bit checkpoints (NVFP4 at 1.32 TiB for an 8xB300 node, MXFP4 at 1.45 TiB for...
DeepSeek released V4 on April 24 and the numbers demand attention. V4-Pro is 1.6 trillion parameters total with 49 billion active, MIT-licensed, native 1M-token context. It scores 80.6% on SWE-bench Verified, putting it within 0.2 points of Claude Opus 4.6. On Terminal-Bench 2...
OpenAI stopped reporting SWE-bench Verified scores. The reason: every frontier model has been trained on the dataset. Morph LLM published the numbers that explain why. Claude Mythos Preview scores 93.9% on the contaminated Verified benchmark. On the new, uncontaminated SWE-ben...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.