Fetching from the wire…
Infra2026-09-03 · source-backed
Leech-lattice vector quantization holds the strongest reported 2-bit quality under its own protocol, but no implementation of the multi-shell decoder existed. This paper supplies one for the full 301-class codebook with a fused dequantize-plus-matvec kernel and measures batch-1 decode GEMV cost. The distinction that matters is that in-VRAM rate is a separate axis from on-disk rate: four bit-exact layouts show binary bit planes beating one-hot masks at 4.80 bits per weight. Run in-process against deployed AWQ (4-bit) and QTIP (2-bit) kernels, the trellis kernel reads 2.40x fewer bytes and runs 2.27x faster, with the time gap tracking the traffic gap.
Each link below shares sources, entities, or timing with this story.
GSQ uses Gumbel-Softmax sampling to match the accuracy of QTIP and AQLM while keeping the deployment simplicity of GPTQ/AWQ. If you're quantizing models for local inference, this eliminates the accuracy-vs-complexity tradeoff.
RAGAS-style evaluation checks correctness against a frozen snapshot, which means routine document updates and corrections can silently break production without moving a dashboard. This ASE 2026 paper defines 11 mutation operators perturbing at both the pre-chunk index level an...
A measurement paper published August 28 ran one adaptive adversary against a seven-layer stack and found failure correlation positive in all fifteen measurable pairs, phi between 0.30 and 0.75, with the joint residual exceeding the multiplicative prediction by up to 0.172. The...
arXiv 2607.26935 argues the human-vs-bot label space can't represent agent traffic: an MLP binary classifier misroutes 39.1% of real agent sessions as human, a SAINT transformer 34.5%, while adding an explicit third class yields agent F1 = 1.000 across all 30 runs. Against a f...
Using op-schema-aware seeded fuzzing against a high-precision fp64 CPU reference on 24 Triton kernels, 15 correct and 9 intentionally buggy, the method caught all 9 buggy variants and passed all 15 controls across five GPU classes (arXiv:2606.20128). Standard kernel benchmarks...
Hugging Face's August 26 post introduces ColBERT-style late-interaction training with MultiVectorEncoder, MultiVectorEncoderTrainer, CachedMultiVectorMultipleNegativesRankingLoss and a MultiVectorInformationRetrievalEvaluator (Hugging Face). Their finetuned mLateOn-medical rea...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.