llama.cpp adds Ling 3.0 VL and tunes Vulkan cooperative matrix for Adreno GPUs
GitHub·low signal
Build b11156 (24 September, #29151) folds Ling 3.0 VL into the BailingMoeV3 architecture, and its commit carries an 'Assisted-by: Scout' AI disclosure. Build b11158 (#29328) enables KHR cooperative matrix support in the Vulkan backend for Qualcomm Adreno GPUs, which matters for on-device inference on Snapdragon phones and laptops. b11157 adds CUDA conv3d with implicit GEMM (#29137).