summaryrefslogtreecommitdiff
path: root/ggml-cuda
diff options
context:
space:
mode:
authorIwan Kawrakow <iwan.kawrakow@gmail.com>2024-06-19 18:23:57 +0200
committerIwan Kawrakow <iwan.kawrakow@gmail.com>2024-06-22 12:02:52 +0300
commit7f968d51b4eb6f403bb7dbc1a5bbf98491ff293b (patch)
tree1f6c25e0449aa4684f8b30be5b05b04611c42b68 /ggml-cuda
parentd08ff0df433ee9dd8643afe1cf501c4154067cd2 (diff)
bitnet(scale in a separate tensor): mul -> scale on Metal
Do the mul -> scale replacement on the fly in the Metal backend. This recovers the PP performace and cuts the TG performance degradation in half.
Diffstat (limited to 'ggml-cuda')
0 files changed, 0 insertions, 0 deletions