Tags: sprayberry-code/llama.cpp
Tags
llama-quant : exclude i32 ffn_gate_tid2eid routing table from quantiz… …ation (ggml-org#25787) DeepSeek-V4's ffn_gate_tid2eid tensor is an i32 token-id -> expert-id index table, not weights. It was never added to the name-based exclusion list alongside ffn_gate_inp.weight, so llama-quantize tries to quantize it and fails since i32 cannot convert to a float type. Fixes ggml-org#25754 Signed-off-by: Yash Raj Pandey <yashpn62@gmail.com>
opencl: load and use `kernel_gemm_moe_q6_k_f32_ns` from bin kernel lib ( ggml-org#25797)
tests : initialize all tensors in test_dsv4_hc to avoid NaNs in senti… …nel tensors (ggml-org#25822) Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
ggml-blas: default hadamard mul_mat to cpu routine (ggml-org#25710) Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
vulkan: Support Q2_0 (ggml-org#25430) * vulkan: Support Q2_0 The backend perf tests for mat-vec-mul weren't very good at first (worse than q2_k), doubling the rows per workgroup made a big difference. * reorder * resolve merge conflict, adjust err threshold for f16->q2_0 set_rows
sycl: fix row calculation when K_QUANTS_PER_ITERATION is 1 (ggml-org#… …25690) * sycl: fix incorrect row calculation when K_QUANTS_PER_ITERATION=1 Signed-off-by: Todd Malsbary <todd.malsbary@intel.com> * sycl: use K_QUANTS_PER_ITERATION for non-reordered Q5_K kernel This is the only Q5_K kernel that was not using KQPI. Signed-off-by: Todd Malsbary <todd.malsbary@intel.com> * sycl: add missing second half processing to reordered q5_k Error found while running GGML_SYCL_PRIORITIZE_DMMV=1 \ build/bin/test-backend-ops test -o MUL_MAT Signed-off-by: Todd Malsbary <todd.malsbary@intel.com> * sycl: fix potential off-by-one error Signed-off-by: Todd Malsbary <todd.malsbary@intel.com> * sycl: fix missing row > nrows check Signed-off-by: Todd Malsbary <todd.malsbary@intel.com> --------- Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>
docs: added a note about using OpenCl with Adreno 810 (ggml-org#25786)
PreviousNext