Fetching from the wire…
Models2026-04-30 · source-backed
The release includes LLMs at 3B/8B/30B, speech models (2B for ASR and translation), plus vision and embeddings. All at 512K context, trained on ~15 trillion tokens. The 8B delivering 32B-class performance at a fraction of the compute cost is the real story here. For cost-sensitive deployments, an Apache-licensed 8B that punches above its weight class is worth benchmarking.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source domain / What happened next
Both cover Apache, ASR; reported by the same outlet (huggingface.co); picks up the Apache thread on 2026-07-26.
Shared entities / Same source domain / Earlier coverage
Both cover Apache, IBM Granite; reported by the same outlet (huggingface.co); earlier Apache coverage from 2026-03-17.
Shared entity: Apache / Same source domain / Shared topic / What happened next
Both cover Apache; reported by the same outlet (huggingface.co); overlapping topics (apache, context).
Shared entity: Apache / Same source domain / Shared topic / Earlier coverage
Both cover Apache; reported by the same outlet (huggingface.co); overlapping topics (apache, context).
Shared entity: Apache / Shared topic / Earlier coverage
Both cover Apache; overlapping topics (apache, class, deployment, performance); earlier Apache coverage from 2026-04-13.
Shared entity: Apache / Same source domain / Shared topic / Earlier coverage
Both cover Apache; reported by the same outlet (huggingface.co); overlapping topics (apache, deployment).
Shared entities / Earlier coverage / Tension
Both cover Apache, LLMs; earlier Apache coverage from 2026-04-01; pushes against this story (vs).
Shared entity: LLMs / Shared topic / What happened next / Tension
Both cover LLMs; overlapping topics (above, cost); picks up the LLMs thread on 2026-07-31.