Fetching from the wire…
Public story · 2026-07-30 · high
The new variant is built for everyday coding and Q&A, not long documents, and its docs page picked up 457 points on Hacker News.
Why now: The half-quota variant follows two weeks after K3's July 16 launch and four days after its open weights release on July 26.
Moonshot shipped Kimi K3-256k, a cheaper flagship variant capped at a quarter of the context, two weeks after K3 launched, per its docs page.
The smaller model runs at roughly half the quota of full K3 while matching its output inside that 256k window, according to the docs page. That's a real tradeoff, not just a smaller number: Moonshot is selling reduced context as the cost lever instead of hiding it behind vague throttling or pricing tiers that never say what you're giving up.
The docs page positions K3-256k for everyday Q&A, code completion, routine feature work, and single-file edits, the kind of jobs that rarely touch more than a few files at once. It takes image input but not video, and offers low, high, and max reasoning effort settings, with high as the default. The page doesn't say what happens once a task needs more than 256k tokens, only that Moonshot built this version for work that doesn't.
Full K3, a 2.8-trillion-parameter model, launched July 16 with open weights following on July 26. The docs page for the smaller variant picked up 457 points on Hacker News, more attention than most model announcements get for what amounts to a context downgrade.
Pricing a smaller context window as an explicit discount reads more honest than the usual frontier playbook of raising limits and calling it progress. Watch whether other labs start selling narrower windows on purpose instead of only wider ones.
Each link below shares sources, entities, or timing with this story.
The open-weight race just changed constraint. Moonshot AI suspended all new consumer subscriptions on July 20, roughly 48 hours after Kimi K3 launched, because request volume pushed its compute cluster to capacity. Remaining GPUs are reserved for existing paid subscribers. Tec...
Moonshot exposes an Anthropic-compatible endpoint, so pointing Claude Code at K3 means setting the Anthropic base URL and supplying a Moonshot key. No new CLI, no config rewrite. Hosted at $3/$15 per Mtok, same tier as Claude Sonnet 4.6, and Artificial Analysis scores K3 at 57...
OSTP Director Michael Kratsios posted July 22 that Moonshot built "a sophisticated internal platform to conduct large scale distillation against U.S. models," switching access methods to avoid detection, and acquired GB300-equipped servers plus GB300 access in Thailand. TechCr...
Moonshot released K3's open weights July 26 with official guidance calling for 64+ accelerators. WASTE (1,366 stars, created July 28) runs it on a 64GB MacBook Pro at 0.45-0.62 tok/s, keeping the 27.28GB trunk resident and streaming experts from NVMe with 3-bit residual vector...
Alongside WASTE, gavamedia/deltafin (603 stars, created July 28) runs full K3 on a single device with an OpenAI-compatible server. Moonshot published K3's open weights July 26-27 at 2.8T parameters; within 48 hours two separate projects appeared whose entire purpose is fitting...
His July 18 video ("Did Kimi K3 really beat Fable?", 12 minutes, ~72K views) interrogates the claims from Moonshot's July 16 release instead of restating the press cycle. The key context: K3's headline win is a single-category result on Arena's Frontend Code eval, sitting alon...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.