ik_llama.cpp.git - Unnamed repository; edit this file 'description' to name the repository.

Age	Commit message (Collapse)	Author
2024-09-02	Do not process prompts containing binary data for escapes (#33)	Kawrakow
	Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
2024-09-01	Zen4 Flash Attention (#32)	Kawrakow
	* Zen4 flash attention: moving useful parts from the kq_fused_softmax branch * Add flash attention with soft-cap and fix D = 256 case * Flash attention refinements * Update FlashAttn comment --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
2024-08-31	Fix build when iqk_mul_mat is disabled (#31)	Kawrakow
	Ref #29 Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
2024-08-27	Faster Gemma2 (#27)	Kawrakow
	* soft_cap_max: initial CPU version of fused softcap + soft_max With this vanilla CPU implementation I'm already getting a ~3% speedup for Gemma-2-9b and a prompt of 8192 tokens. * soft_cap_max: WIP - something is wrong with CUDA * soft_cap_max: looks good on CPU and CUDA * Add softcap to flash attention Just CPU and CUDA for now (but, as we know, flash attention on the CPU is useless in llama.cpp). On CUDA this improves PP performance quite a bit, especially for long contexts. E.g., for PP-16384, I now get 3777 t/s. Without this change, one cannot use FA, and one gets 2300 t/s (after fusing softcap and softmax), or 2000 t/s without the fused softcap+softmax. In comparison, mainline llama.cpp has PP-16384 = 1549 t/s before PR-8542 (where Johannes Gaessler has also added softcap to FA), and PP-16384 = 3097 t/s after this PR. * soft_cap_max: Metal * Flash attention with softcap: Metal --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
2024-08-21	softcap: minor improvement (#24)	Kawrakow
	Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
2024-08-20	Fused soft cap and SIMD-ified GeLU (#9)	Kawrakow
	* Softcap: WIP Fuses scale + tanh + scale as used for softcaping in some models. Just CPU for now. ~1.4% for PP-512 on Gemma2-9b, no effect on TG. Somewhat surprisingly the improvement does not increase as I go to longer contexts. Gemma2 does softcap on KQ, which grows quadratically with context length, so I would have thought the benefit from fusing scale, tanh, scale would increase. But no, no luck. softcap: CUDA * softcap: CUDA ~1% speedup for Gemma2-9b * softcap: Metal and NEON About 1% speedup. * Simdified gelu Gives ~1% speedup for Gemma2-9b prompt processing on AVX512/AVX2. It looks like the gelu operation is memory bound on my CPU's after SIMD-ifying it. By not using the 128 kb gelu lookup table we gain a small advantage. On the M2-Max the lookup table is slightly faster than the SIMD version, so left the lookup table for ARM_NEON. * softcap, tanh: avoid NaNs for large arguments (AVX2, AVX512) Not that I have encountered this in practice, but just to be sure. This does it for AVX512 and AVX2, still need a guard for ARM_NEON. * llama-bench: add ability to turn off warmup runs So we don't need to wait forever on, e.g., benchmarks involving long contexts. * softcap, tanh: avoid NaNs for large arguments (NEON) --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
2024-08-20	iq4_k: use iq5_k also when n_gqa = 2 (#23)	Kawrakow
	This improves size vs quality balance for Gemma-2 models. Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
2024-08-19	AVX2 quantization for Q8_K (#22)	Kawrakow
	It has been there for a while, but forgot to add here. Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
2024-08-19	quantize_stats: print rmse and max error as fraction of <x> (#21)	Kawrakow
	This allows for a better comparison between different models or different tensors of the same model where the magnitude of the model weights may differ. Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
2024-08-19	iq2_k: slightly better bpw - accuracy compromise (#20)	Kawrakow
	For LLaMA-3.1 models: * It is better to quantize all of attn_v with iq3_k instead of half of attn_v with iq4_k * Quantizing attn_output with iq3_k results in a larger PPL decrease compared to what one expects from the added bpw. Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
2024-08-14	Skip barriers of noops (#19)	Kawrakow
	GGML_OP_RESHAPE, GGML_OP_VIEW, GGML_OP_PERMUTE, GGML_OP_TRANSPOSE, along with GGML_OP_NONE, are all noops. I.e., nothinh happens. But ggml still has a barrier after them, which wastes time. The waste is not too bad for large models where computations are long compared to the time taken for thread synchronization. But for small models skipping those unnecessary waits makes a significant difference. E.g., for the 99M TriLMamodel, TG-500 goes up to 1426 t/s from 1240 t/s. Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
2024-08-12	Update README.md	Kawrakow

2024-08-12	Merge mainline - Aug 12 2024 (#17)	Kawrakow
	* Merge mainline * Fix after merge * Remove CI check --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
2024-08-09	Fix Makefile	Iwan Kawrakow
	I always use cmake, so had forgotten to pay attention to the Makefile.
2024-08-09	Fix Zen4 implementation of iq3_k, iq4_k, iq5_k	Iwan Kawrakow
	See comments in f3a823ce729a7db33e7d4375eae7291bbe6196db
2024-08-09	iq6_k: AVX2	Iwan Kawrakow

2024-08-09	iq6_k: Metal	Iwan Kawrakow
	About 4% slower than Q6_K for PP-512, but 10% faster for TG-128. Someone has screwed up Q6_K TG performance on Metal? With the cobntinuous "improvements" in ggml I wouldn't be surprised. Need to look into it later.
2024-08-09	iq6_k: NEON	Iwan Kawrakow
	Respectable performance, only slightly slower than Q6_K.
2024-08-09	iq6_k: slightly better Zen4 iqk_mul_mat	Iwan Kawrakow
	We now arrive at pp-512 = 147 t/s for LLaMA-3.1-8B. TG-128 is 9.5 t/s. This is better than last commit, but still kind of slow compared to Q6_K. My last commit message is wrong: also iq3_k needs a fix for overflow.
2024-08-09	iq6_k: Zen4 iqk_mul_mat	Iwan Kawrakow
	We need to do 4 shuffles to get the non-uniform values, so this makes it slower than other iqX_k quants. And then I realized that I was using the standard Zen4 template for all iqX_k quants. The standard template converts the 32-bit integers obtained after _mm512_dpbusds_epi32 back to 16 bits, and then multiples with 16-bit block scales. But this can overfow for iq4_k, iq5_k, and iq6_k. I guess, I did not notice with iq4_k and iq5_k because the PPL difference to CUDA was relatively small, and I attributed it to Q8_K not being accurate enough for the activations. But for iq6_k the PPL difference was much too big to be attributable to Q8_K inaccuracies, so that's when I realized that I cannot be packing the _mm512_dpbusds_epi32 result into 16 bit for 4-,5-,6-bit iqX_k quants. For now I fixed it for iq6_k, but the outcome is that it is significantly slower than Q6_K: I get PP-512 = 125 t/s for LLaMA-3.1-8B vs 180 t/s for Q6_K, so I need to look for a better approach.
2024-08-09	iq6_k: CUDA dot product	Iwan Kawrakow
	90.2 t/s for LLaMA-3.1-8B. Q6_K gives 91.2 t/s, so we are good.
2024-08-09	iq6_k: CUDA dequantize	Iwan Kawrakow
	We get a slightly better PPL for LLaMA-3.1-8B compared to q6_K (0.14% vs 0.26% quantization error).
2024-08-09	iq6_k: WIP (quantize/dequantize)	Iwan Kawrakow

2024-08-09	iq6_k: WIP (nothing works)	Iwan Kawrakow

2024-08-07	Adding IQ2_TN for use with ternary models (#13)	Kawrakow
	* iq2_tn: TriLM specific 2.0625 bpw quantization Quantize/dequantize/scale dot product. I get 46 t/s for the TriLM-3.9B with any SIMD! Finally a compiler doing a decent job auto-vectorizing the scalar implementation. * iq2_tn: AVX512 Just reusing the k-quants template gets us to PP-512 = 376 t/s, TG-128 = 47.6 t/s for TriLM-3.9B. * iq2_tn: AVX512 With this tweak we get to PP-512 = 431 t/s. * iq2_tn: AVX512 With this tweak we get TG-128 = 19.58 / 35.18 t/s for 1 / 2 threads. At 4 threads we saturate at 48.41 t/s, and then performance slowly degrades with increasing number of threads. * iq2_tn: AVX2 PP512 = 440 t/s on the Ryzen-5975WX. We should be able to do better. * iq2_tn: initial NEON version * iq2_tn: NEON For TriLM-3.9B running on the M2-Max we get PP-512 = 193.5 t/s, TG-128 = 75.5 t/s. This is in line with what we have for iq2_bn ant 3.3B Bitnet. * iq2_tn: Metal For TriLM-3.9B on a 30-core M2-Max we get PP-512 = 890 t/s, TG-128 = 98.5 t/s. * iq2_tn: CUDA For TriLM-3.9B running on RTX-4080 we get PP-512 = 9936 t/s, TG-128 = 299.2 t/s. * iq2_tn: AVX2 PP improvement We now get PP-512 = 490.73 t/s for TriLM-3.9B on the Ryzen-5975WX. We have PP-512 = 636.61 t/s for Bintnet-3B quantized with iq2_bn. Bintnet-3B is actually 3.4B, TriLM-3.9B is 3.99B, so we would expect 3.43/3.99 * 636 = 546 t/s, so it seems we still have something that is not quite optimal in iq2_tn. * iq2_tn: small NEON improvement For TriLM-3.9B we now get PP-512 = 206.6 t/s and TG-128 = 76.4 t/s. --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
2024-08-05	q2_K: allow it to detect ternary nets and quantize accordingly	Iwan Kawrakow

2024-08-05	Update README.md	Kawrakow
	There have been a few minor improvements here and there, so updated the AVX2 Bitnet performance values to current main branch.
2024-08-05	iq3_k, iq5_k: faster quantization	Iwan Kawrakow
	Just use the same trick as iq4_k
2024-08-03	iq4_k: speedup quantization by a factor of ~2	Iwan Kawrakow

2024-08-01	Add copyright notice	Iwan Kawrakow

2024-08-01	iq2/3_k: tiny bit faster Metal dot products	Iwan Kawrakow

2024-08-01	iq3_k: slightly faster Metal dequantize kernel	Iwan Kawrakow
	PP-512 goes to 473 t/s up from 452 t/s.
2024-08-01	iq3_k: Metal dot product	Iwan Kawrakow
	Quite slow: 43 t/s for a 7B model
2024-08-01	iq2_k: Metal dot product finally works	Iwan Kawrakow
	It is slow: 45.4 t/s for 7B model vs 50 t/s for iq2_xs, or 63.3 t/s for q2_K_S.
2024-08-01	iq3_k: Metal dequantize	Iwan Kawrakow

2024-08-01	iq3_k: NEON	Iwan Kawrakow

2024-08-01	iq3_k: AVX2 iqk_mul_mat	Iwan Kawrakow
	We get PP-512 = 196 t/s for LLaMA-3.1-8B on the Ryzen-5975WX.
2024-08-01	iq3_k: AVX512 iqk_mul_mat	Iwan Kawrakow
	We get PP-512 = 180 t/s, TG-128(4 threads) = 16.35 on the Ryzen-7950X for LLaMA-3.1-8B. In comparison, iq3_s has PP-512 = 96 t/s, TG-128 = 7.6 t/s with iqk_mul_mat, and PP-512 = 28 t/s, TG-128 = 6.8 t/s in mainline llama.cpp
2024-08-01	iq3_k: faster CUDA dot product	Iwan Kawrakow
	138 t/s for LLaMA-3.1-8B, which is almost on par with iq3_s.
2024-08-01	iq3_k: CUDA dot product	Iwan Kawrakow
	Slightly slower than iq3_s - 132 t/s vs 138 t/s for LLaMA-3.1-8B.
2024-08-01	iq3_k: Basics	Iwan Kawrakow
	Quantize/dequantize, CUDA dequantize. PPL of LLaMA-3.1-8B is better than iq3_s and iq3_m.
2024-08-01	iq2_k: very slightly better CUDA dot product	Iwan Kawrakow
	169.2 t/s vs 167.8 t/s before.
2024-08-01	iq2_k: better CUDA dot product	Iwan Kawrakow
	Almost on par with iq2_xs (168 t/s vs 172 t/s).
2024-08-01	iq2_k: CUDA dot product finally works	Iwan Kawrakow
	Performance is pathetic: 140 t/s for LLaMA-3.1-8B vs 172 t/s for iq2_xs.
2024-08-01	iq5_k: CUDA dot product finally works	Iwan Kawrakow

2024-08-01	Factor out iqk CUDA dot products	Iwan Kawrakow
	I cannot possibly wait for a 5 minutes nvcc compilation each time I touch vecdotq.cuh. Also, cmake was adding --options-file X.rsp to the nvcc compile commands, which confuses clangd, so I have turned that off.
2024-08-01	iq5_k: CUDA dot product still not working	Iwan Kawrakow

2024-08-01	iq5_k: Metal	Iwan Kawrakow
	Performance is roughly on par with q5_0.
2024-08-01	iq5_k: NEON	Iwan Kawrakow

2024-08-01	iq5_k: AVX512	Iwan Kawrakow