Sampled Layerwise Proofs Verify Llama-2-70B Inference on a CPU Host and Batch Twelve Requests Into One Proof at 6.5x Lower Cost
arXiv·low signal
arXiv 2609.27367 (23 Sep) commits the boundary activations of every chunk of an inference trace before any challenge, then proves a verifier-chosen subset. On TinyLlama, proving 7 of 47 chunks takes 22% of the time of proving all of them. Packing twelve requests into one trace costs 6.5x less than twelve separate proofs, and a Llama-2-70B run produced a 4.34 MiB proof in 1,259 s that verified in 46.3 s without the weights.