Markets
BaseRT Claims 6.4x Faster Than llama.cpp and 3.9x Faster Than MLX on Apple Silicon
BaseRT launched July 19 (open source, Apple silicon) claiming 6.4x throughput over llama.cpp and 3.9x over MLX — the two default local-inference runtimes for Mac developers. If the numbers hold under independent benchmarking, the practical effect is that a class of workloads currently paying per-token to a hosted API becomes economical on a laptop, which is the actual mechanism by which inference SaaS gets disintermediated. Treat the claims as unverified until third parties reproduce them; single-source vendor benchmarks in this category have historically been measured under favorable batch and quantization settings.
Source
↳ Follow the thread