Fetching from the wire…
Public story · 2026-07-22 · high
One tier setting replaces separate per-provider priority flags for every model behind Vercel's Gateway.
Why now: The tier addition sits in the same changelog that added Poolside's Laguna S 2.1 the same week, per the July 22 briefing.
Vercel added service tiers to its AI Gateway: a faster tier for interactive requests, a cheaper one for batch jobs, per its changelog. For teams routing across model providers, that's one less integration to maintain, since the same tier setting applies no matter which model answers the request.
Priority-versus-flex pricing already exists at the provider level; several model providers run their own version of it. What's new is putting one tier abstraction in front of all of them. A developer sets a single config instead of learning each provider's separate priority system.
A chat interface where someone's waiting gets the fast tier. A nightly batch job summarizing documents doesn't need it, so it runs flex and cheaper.
Vercel also added Poolside's Laguna S 2.1 the same week, in a free 256K-context version and a paid 1M-context version.
Vercel's changelog doesn't publish the latency or cost gap between tiers, and I haven't tested it myself. Worth watching whether other gateway products answer with their own uniform tier layer.
Each link below shares sources, entities, or timing with this story.
Satya Nadella said companies routing everything through a single proprietary lab may not survive. His argument: you hand that lab your most sensitive business context, and the lab can turn it against you as a competitor. His prescription is an orchestration layer — keep the ha...
The beta lets you set speed to fast and the gateway serves the fast tier wherever a model offers one, using one request shape for every provider. Small thing, real annoyance removed: each lab exposes latency tiers differently, so today you either hardcode per-provider flags or...
Routed through minimax/minimax-m3-free and minimax/minimax-m2.7-free via GMI Cloud. The free IDs stop returning after the window. Concretely useful for benchmarking MiniMax against your current default without standing up an account. ---
Vercel CEO Guillermo Rauch announced open-source, bring-your-own-model templates for both v0 and Vercel Agent. Powered by the AI SDK, Vercel AI Gateway, and Sandbox. The template supports Claude Code, OpenAI Codex CLI, GitHub Copilot CLI, Cursor CLI, Gemini CLI, and opencode....
Announced August 7: Hermes can use AI Gateway as its inference layer for 200+ models with no token markup and per-request dashboard visibility, and execute shell commands inside an isolated Vercel Sandbox microVM instead of on your machine, with Node.js 24/22 and Python 3.13 a...
Announced at Vercel Ship 2026, AI SDK 7 converts a model-abstraction library into a toolkit for agents that reason, call tools, run across many turns, and work across files and sandboxes, behind one provider-agnostic interface. Vercel paired it with Eve (define agent instructi...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.