ik_llama.cpp.git - Unnamed repository; edit this file 'description' to name the repository.

diff options

author	Kawrakow <iwankawrakow@gmail.com>	2024-12-11 14:20:27 +0100
committer	GitHub <noreply@github.com>	2024-12-11 14:20:27 +0100
commit	9469af87f71a8d174064e142ae8eff390d387c6c (patch)
tree	1e6aad8a09d817d6ea57da2ae1ef84f16b416f13 /include/llama.h
parent	e0adb8b1227dd4622a91a5b3680b4af2e36d32f4 (diff)

Better ARM_NEON implementation for R4 quants (#135)

* q6_k_r4: Better ARM implementation PP-512(LLaMA-3.1-8B) is now 104.2 t/s up from 83.2 t/s. I.e., q6_k_r4 now beats q6_0_r4. * q5_k_r4: Better ARM implementation PP-512(LLaMA-3.1-8B) is now 107.8 t/s up from 96.9 t/s. I.e., q5_k_r4 now beats q5_0_r4. * q4_k_r4: Better ARM implementation PP-512(LLaMA-3.1-8B) is now 122.1 t/s up from 110 t/s. I.e., q4_k_r4 is now (nearly) on par with q4_0_r4. * iq4_xs_r4: Better ARM implementation PP-512(LLaMA-3.1-8B) is now 131.3 t/s up from 115.8 t/s. iq4_xs_r4 is now the prompt processing champion on ARM. --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>

Diffstat (limited to 'include/llama.h')

0 files changed, 0 insertions, 0 deletions


context:
space:
mode: