llama : remove LLAMA_MAX_DEVICES and LLAMA_SUPPORTS_GPU_OFFLOAD (#5240)

* llama : remove LLAMA_MAX_DEVICES from llama.h ggml-ci * Update llama.cpp Co-authored-by: slaren <slarengh@gmail.com> * server : remove LLAMA_MAX_DEVICES ggml-ci * llama : remove LLAMA_SUPPORTS_GPU_OFFLOAD ggml-ci * train : remove LLAMA_SUPPORTS_GPU_OFFLOAD * readme : add deprecation notice * readme : change deprecation notice to "remove" and fix url * llama : remove gpu includes from llama.h ggml-ci --------- Co-authored-by: slaren <slarengh@gmail.com>
author: Georgi Gerganov <ggerganov@gmail.com> 2024-01-31 17:30:17 +0200
committer: GitHub <noreply@github.com> 2024-01-31 17:30:17 +0200
commit: 5cb04dbc16d1da38c8fdcc0111b40e67d00dd1c3 (patch)
tree: 3ef8dc640d5c08466309c09a8ac2963bb760af06 /README.md
parent: efb7bdbbd061d087c788598b97992c653f992ddd (diff)
1 files changed, 2 insertions, 1 deletions
diff --git a/README.md b/README.md
index 7746cb51..e6ed1d42 100644
--- a/README.md
+++ b/README.md
@@ -10,7 +10,8 @@ Inference of [LLaMA](https://arxiv.org/abs/2302.13971) model in pure C/C++
 
 ### Hot topics
 
-- ⚠️ Incoming backends: https://github.com/ggerganov/llama.cpp/discussions/5138
+- Remove LLAMA_MAX_DEVICES and LLAMA_SUPPORTS_GPU_OFFLOAD: https://github.com/ggerganov/llama.cpp/pull/5240
+- Incoming backends: https://github.com/ggerganov/llama.cpp/discussions/5138
   - [SYCL backend](README-sycl.md) is ready (1/28/2024), support Linux/Windows in Intel GPUs (iGPU, Arc/Flex/Max series)
 - New SOTA quantized models, including pure 2-bits: https://huggingface.co/ikawrakow
 - Collecting Apple Silicon performance stats:
author	Georgi Gerganov <ggerganov@gmail.com>	2024-01-31 17:30:17 +0200
committer	GitHub <noreply@github.com>	2024-01-31 17:30:17 +0200
commit	5cb04dbc16d1da38c8fdcc0111b40e67d00dd1c3 (patch)
tree	3ef8dc640d5c08466309c09a8ac2963bb760af06 /README.md
parent	efb7bdbbd061d087c788598b97992c653f992ddd (diff)