diff options
author | Kawrakow <iwankawrakow@gmail.com> | 2024-12-11 14:20:27 +0100 |
---|---|---|
committer | GitHub <noreply@github.com> | 2024-12-11 14:20:27 +0100 |
commit | 9469af87f71a8d174064e142ae8eff390d387c6c (patch) | |
tree | 1e6aad8a09d817d6ea57da2ae1ef84f16b416f13 /examples | |
parent | e0adb8b1227dd4622a91a5b3680b4af2e36d32f4 (diff) |
Better ARM_NEON implementation for R4 quants (#135)
* q6_k_r4: Better ARM implementation
PP-512(LLaMA-3.1-8B) is now 104.2 t/s up from 83.2 t/s.
I.e., q6_k_r4 now beats q6_0_r4.
* q5_k_r4: Better ARM implementation
PP-512(LLaMA-3.1-8B) is now 107.8 t/s up from 96.9 t/s.
I.e., q5_k_r4 now beats q5_0_r4.
* q4_k_r4: Better ARM implementation
PP-512(LLaMA-3.1-8B) is now 122.1 t/s up from 110 t/s.
I.e., q4_k_r4 is now (nearly) on par with q4_0_r4.
* iq4_xs_r4: Better ARM implementation
PP-512(LLaMA-3.1-8B) is now 131.3 t/s up from 115.8 t/s.
iq4_xs_r4 is now the prompt processing champion on ARM.
---------
Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
Diffstat (limited to 'examples')
0 files changed, 0 insertions, 0 deletions