Fetching from the wire…
Top 5 · 2026-08-06 · source-backed
$1.25 input / $4.25 output per million tokens for the standard tier. $0.10 / $0.20 for the Contributor tier, where Meta trains on your usage and feedback.
That's roughly 12x on input, 21x on output, and it's the clearest number anyone has published on what your proprietary source code is worth as training data. Meta launched Muse Code in beta on August 5 alongside Muse Spark 1.2, with MacRumors and others reporting the pricing tiers.
Every procurement conversation about coding agents has had this question floating around it unpriced. "Are they training on our code?" gets a policy answer, a DPA, a checkbox. Meta just put a dollar figure on it and made it a menu item. Whatever you think of that, it's more honest than the alternative, and it means the next time someone on your team argues for the cheap tier you can say precisely what it costs: 92% off, paid in source code.
The product itself is genuinely interesting. Muse Code keeps specialized background sub-agents alive across a session so context accumulates rather than resetting each turn, and spawns parallel sub-agents in isolated worktrees for large tasks. It bundles /plan, /grill and /goal skills plus local event logging for crash recovery. Meta says Muse Spark 1.2 came from significantly scaled coding-task training compute and more diverse training environments. The r/singularity thread hit 221 upvotes.
Now the skepticism. Meta benchmarks against Terminal-Bench 2.1 and DeepSWE 1.1 in the announcement and publishes no numeric scores. None. A release post that names its benchmarks and omits its results is telling you something. Press-reported contributor-tier figures also don't fully agree with each other across outlets, so treat the exact numbers as directionally right rather than quotable to the cent.
Willison ran his standard pelican-on-a-bicycle SVG test and called 1.2 "a small but material improvement" over 1.1. His actual verdict on the release is the line worth stealing: "the most important characteristic of any model these days is long-sequence agentic tool calling." Not single-turn reasoning quality, not the leaderboard row. How many tool calls deep it stays coherent.
That's a better evaluation heuristic than anything in the benchmark table, and it's testable on your own workload in an afternoon. Take a real multi-step task from your repo, run it under each candidate model, and count where coherence breaks. Fifteen tool calls? Forty? That number will predict your day-to-day experience better than any published score.
Worth pairing with the Kilo Code data point: co-founder Emilie Schario says her engineers now read or write code directly about 1% of the time, and her cost playbook is frontier models for architecture, open-weight models for everything else, adopted after customers told her "I accidentally spent my whole AI budget for the year." Replit's Amol Jain runs the more conservative posture, "human on the loop, not human in the loop," where an agent risk-scores every PR and only low-risk ones self-merge.
Each link below shares sources, entities, or timing with this story.
Kilo Code supports JetBrains / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Kilo Code supports JetBrains); both cover Bench, Kilo Code, Terminal; overlapping topics (agent, code, model).
Meta partners with Google / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Meta partners with Google); both cover MacRumors, Meta, Whatever; reported by the same outlet (macrumors.com).
SaaStr uses Replit / Shared entity: Replit / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (SaaStr uses Replit); both cover Replit; overlapping topics (agent, call, cost, human, model).
Meta uses Gemini / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Meta uses Gemini); both cover Bench, Terminal; overlapping topics (agent, code, coding, tier).
Kilo Code uses OpenRouter / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Kilo Code uses OpenRouter); both cover Bench, Terminal; overlapping topics (coding, cost, model, number).
Meta uses Gemini / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Meta uses Gemini); both cover Bench, Terminal; overlapping topics (agent, coding, model).
Meta partners with Google / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Meta partners with Google); both cover Bench, Terminal; overlapping topics (agent, coding, model).
Meta released Llama / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Meta released Llama); both cover Meta, Muse Spark; overlapping topics (benchmark, meta, model).