Voices
Kwindla Hultman Kramer open-weights PhoneLLM, a 30B/3.5B-active voice-agent model claiming GPT 5.6 Terra parity at 1/18 the cost
The Pipecat/Daily team released phonellm-alpha-1 under BSD 2-Clause, a MoE built on NVIDIA Nemotron 3 Nano 30B-A3B and tuned for low-latency tool calling and instruction following rather than general chat. The model card claims parity with GPT 5.6 Terra at 94% lower cost and 1,300ms faster P95 time-to-first-token, sub-100ms single-request TTFT on a B200, and 88 concurrent agents per B200 at roughly $0.00025 per agent-minute. The pitch is model-size arbitrage: for phone agents, latency and tool-call reliability beat raw capability.
↳ Follow the thread