Sources
Cerebras is serving Qwen 3.8 27B at roughly 1,500 tokens/second on its public pay-as-you-go endpoints
Cerebras added qwen-3.8-27b to its public model catalog with 64k context on the free tier and 128k on paid, at approximately 1,500 tokens/second, sitting alongside gpt-oss-120b at approximately 3,000 tokens/second. Those are the only two models on the public endpoints, with everything else pushed to Dedicated Endpoints. At that throughput a 27B open model becomes usable for interactive agent loops where the bottleneck is round-trip latency rather than model quality, which is the practical reason to care about a mid-size model on a wafer-scale part.
↳ Follow the thread