MicroQonv Cuts Microscaling Activation Quantization Cost Up to 9x in Conv Layers by Quantizing Before im2col
arXiv·low signal
arXiv 2609.28358 (23 Sep) quantizes each tensor once and applies a channel-batch-first im2col after quantization, where naive pipelines quantize twice and inflate activations first. It reduces quantization cost 2x for weights and gradients and up to 9x for activations, cuts memory movement up to 7.53x against full precision, and reduces MX activation traffic 3.5x on YOLOv8-nano, with negligible accuracy loss.