Sitelet https://github.com/sprayberry-code/llama.cpp/tags
Skip to content

Tags: sprayberry-code/llama.cpp

Tags

b10067

Toggle b10067's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
llama-quant : exclude i32 ffn_gate_tid2eid routing table from quantiz…

…ation (ggml-org#25787)

DeepSeek-V4's ffn_gate_tid2eid tensor is an i32 token-id -> expert-id
index table, not weights. It was never added to the name-based
exclusion list alongside ffn_gate_inp.weight, so llama-quantize tries
to quantize it and fails since i32 cannot convert to a float type.

Fixes ggml-org#25754

Signed-off-by: Yash Raj Pandey <yashpn62@gmail.com>

b10066

Toggle b10066's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
opencl: load and use `kernel_gemm_moe_q6_k_f32_ns` from bin kernel lib (

ggml-org#25797)

b10064

Toggle b10064's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
opencl: transpose q4_K noshuffle scales for coalesced reads (ggml-org…

…#25805)

b10063

Toggle b10063's commit message
sync : ggml

b10061

Toggle b10061's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
tests : initialize all tensors in test_dsv4_hc to avoid NaNs in senti…

…nel tensors (ggml-org#25822)

Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>

b10059

Toggle b10059's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
ggml-blas: default hadamard mul_mat to cpu routine (ggml-org#25710)

Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>

b10058

Toggle b10058's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
vulkan: Support Q2_0 (ggml-org#25430)

* vulkan: Support Q2_0

The backend perf tests for mat-vec-mul weren't very good at first (worse than
q2_k), doubling the rows per workgroup made a big difference.

* reorder

* resolve merge conflict, adjust err threshold for f16->q2_0 set_rows

b10057

Toggle b10057's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
sycl: fix row calculation when K_QUANTS_PER_ITERATION is 1 (ggml-org#…

…25690)

* sycl: fix incorrect row calculation when K_QUANTS_PER_ITERATION=1

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* sycl: use K_QUANTS_PER_ITERATION for non-reordered Q5_K kernel

This is the only Q5_K kernel that was not using KQPI.

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* sycl: add missing second half processing to reordered q5_k

Error found while running

  GGML_SYCL_PRIORITIZE_DMMV=1 \
  build/bin/test-backend-ops test -o MUL_MAT

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* sycl: fix potential off-by-one error

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

* sycl: fix missing row > nrows check

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

---------

Signed-off-by: Todd Malsbary <todd.malsbary@intel.com>

b10056

Toggle b10056's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
opencl: add ABS op (ggml-org#25115)

b10054

Toggle b10054's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
docs: added a note about using OpenCl with Adreno 810 (ggml-org#25786)