Fetching from the wire…
Models2026-09-24 · source-backed
Separate state and action encoders trained with an InfoNCE objective, scoring candidate actions by embedding alignment. Pre-trained on 60M Nemotron Q&A pairs, post-trained on 1M agentic trajectories. The team reports Jev-level zero-shot results on computer use, games and tool calling, and after fine-tuning, 81.6% on DeepSWE and 87.6% on Terminal-Bench 2.1 (Contrastive Language Models). Action embeddings cache, so latency claims run 4-13x better. The coding numbers need independent replication before anyone builds on them.
Each link below shares sources, entities, or timing with this story.
NVIDIA claims up to 4x faster output and 30% faster agentic task completion versus comparable models, aimed explicitly at high-volume narrow work: code review, tool use, security monitoring, billing triage (NVIDIA). They also published Nemotron-RL-Agentic-Terminal-Pivot, an ag...
NVIDIA debuted the Nemotron 3 family for agentic AI. Nano (30B params, 3B active via MoE) delivers 78% HumanEval, native 1M-token context, and 4x higher throughput than Nemotron 2 via hybrid Mamba-Transformer architecture. Critical developer angle: NVIDIA open-sources NeMo Gym...
RTK has almost 80,000 GitHub stars and a simple promise. It sits between your coding agent and the shell, trims noisy command output before the model reads it, and claims 60-90% savings. Quesma ran it on Terminal-Bench 2.1 and found costs went up. With RTK on, average cost per...
Nemotron-Terminal (arXiv:2602.21193, 66 HF upvotes) — First systematic study of data engineering for terminal/CLI agents. Terminal-Task-Gen pipeline with Dockerized environment interaction. Qwen3-initialized 8B model goes from 2.5% to 13.0% on Terminal-Bench 2.0. All checkpoin...
The comparison is against GB300 NVL72, with 35x lower cost per million tokens, measured on the SemiAnalysis AgentX benchmark using real recorded agentic coding sessions with context growth, tool calls and sub-agent spawning preserved (NVIDIA). DeepSeek V4 Pro and Qwen3.5 were...
The Series C, led by Aramco Ventures with NVIDIA, Vista, and others, funds a plan to grow capacity roughly 50x over five years, serving DeepSeek, Nemotron, MiniMax, and Kimi. Capital is still flowing hard into open-model serving infrastructure, which is the supply side of the...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.