summaryrefslogtreecommitdiff
path: root/convert-hf-to-gguf.py
diff options
context:
space:
mode:
authorIwan Kawrakow <iwan.kawrakow@gmail.com>2024-06-17 18:51:00 +0200
committerIwan Kawrakow <iwan.kawrakow@gmail.com>2024-06-22 12:02:52 +0300
commit1de6476d751a02978b035feb38066462c4382877 (patch)
tree56957851a3b91eb09100d0284bbda8e976723ade /convert-hf-to-gguf.py
parentf97a3296387c3fd0ac96324263cecc3d84edfd43 (diff)
bitnet 2 bpw: NEON implementation
We get PP-512 = 190 t/s and TG-128 = 75 t/s. 2 bpw TG on the CPU beats 1.75 bpw on the GPU!
Diffstat (limited to 'convert-hf-to-gguf.py')
0 files changed, 0 insertions, 0 deletions