ik_llama.cpp.git - Unnamed repository; edit this file 'description' to name the repository.

Age	Commit message (Collapse)	Author
2023-12-29	clip : use ggml_backend_buffer_is_host (#4205)	Georgi Gerganov

2023-12-29	clip : enable gpu backend (#4205)	Steward Garcia
	* clip: enable CUDA backend * add missing kernels * add enough padding for alignment * remove ggml_repeat of clip.cpp * add metal backend * llava : fixes - avoid ggml_repeat - use GGML_USE_ instead of CLIP_USE_ macros - remove unused vars --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2023-12-29	cmake : fix ld warning duplicate libraries libllama.a (#4671)	Cuong Trinh Manh
	* fix "ld: warning: ignoring duplicate libraries: '../libllama.a'" * fix warning in example.
2023-12-29	llava-cli : refactor to use sampling library (#4669)	Justine Tunney
	This change makes it possible to use flags like `--grammar` when using the `llava-cli` program. The rest is just code cleanup deleting a long standing TODO comment. This change also ensures that logging information is emitted to stderr which helps the `llava-cli` command be more friendly to shell scripts. See Mozilla-Ocho/llamafile@1cd334f
2023-12-29	server : replace sleep with condition variables (#4673)	Justine Tunney
	The server currently schedules tasks using a sleep(5ms) busy loop. This adds unnecessary latency since most sleep implementations do a round up to the system scheduling quantum (usually 10ms). Other libc sleep impls spin for smaller time intervals which results in the server's busy loop consuming all available cpu. Having the explicit notify() / wait() code also helps aid in the readability of the server code. See mozilla-Ocho/llamafile@711344b
2023-12-29	server : fix OpenAI server sampling w.r.t. penalty. (#4675)	SakuraUmi

2023-12-29	server : allow to generate multimodal embeddings (#4681)	Karthik Sethuraman

2023-12-29	main-cmake-pkg : fix build issue (#4665)	andrijdavid
	* Fix main-cmake-pkg compilation * Use glob to load common files * cmake : fix trailing whitespace --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2023-12-29	llama.swiftui : fix infinite loop, ouput timings, buff UI (#4674)	Peter Sugihara
	* fix infinite loop * slight UI simplification, clearer UX * clearer UI text, add timings to completion log
2023-12-28	Fix OpenAI server sampling w.r.t. temp and seed (#4668)	Justine Tunney
	The default values for tfs_z and typical_p were being set to zero, which caused the token candidates array to get shrunk down to one element thus preventing any sampling. Note this only applies to OpenAI API compatible HTTP server requests. The solution is to use the default values that OpenAI documents, as well as ensuring we use the llama.cpp defaults for the rest. I've tested this change still ensures deterministic output by default. If a "temperature" greater than 0 is explicitly passed, then output is unique each time. If "seed" is specified in addition to "temperature" then the output becomes deterministic once more. See mozilla-Ocho/llamafile#117 See mozilla-Ocho/llamafile@9e4bf29
2023-12-27	finetune : fix output formatting in print_params (#4653)	Daniel Bevenius
	This commit fixes the output formatting in the print_params function which currently looks like this: ```console print_params: n_vocab: 32000 print_params: n_ctx: 128 print_params: n_embd: 4096 print_params: n_ff: 11008 print_params: n_head: 32 print_params: n_head_kv: 32 print_params: n_layer: 32 print_params: norm_rms_eps : 0.000010 print_params: rope_freq_base : 10000.000000 print_params: rope_freq_scale : 1.000000 ``` With this comit the output will look like this: ```console print_params: n_vocab : 32000 print_params: n_ctx : 128 print_params: n_embd : 4096 print_params: n_ff : 11008 print_params: n_head : 32 print_params: n_head_kv : 32 print_params: n_layer : 32 print_params: norm_rms_eps : 0.000010 print_params: rope_freq_base : 10000.000000 print_params: rope_freq_scale : 1.000000 ``` Signed-off-by: Daniel Bevenius <daniel.bevenius@gmail.com>
2023-12-23	server : allow to specify custom prompt for penalty calculation (#3727)	Alexey Parfenov

2023-12-22	lookup : add prompt lookup decoding example (#4484)	LeonEricsson
	* initial commit, going through initializations * main loop finished, starting to debug * BUG: generates gibberish/repeating tokens after a while * kv_cache management * Added colors to distinguish drafted tokens (--color). Updated README * lookup : fix token positions in the draft batch * lookup : use n_draft from CLI params * lookup : final touches --------- Co-authored-by: Leon Ericsson <leon.ericsson@icloud.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2023-12-21	ggml : change ggml_scale to take a float instead of tensor (#4573)	Georgi Gerganov
	* ggml : change ggml_scale to take a float instead of tensor * ggml : fix CPU implementation * tests : fix test-grad0 ggml-ci
2023-12-21	gguf : simplify example dependencies	Georgi Gerganov

2023-12-18	llama.swiftui : add tinyllama 1.1B F16	Georgi Gerganov

2023-12-18	llama.swiftui : add more models	Georgi Gerganov

2023-12-17	llama.swiftui : add bench functionality (#4483)	Georgi Gerganov
	* llama.swiftui : add bench button * llama.swiftui : initial bench functionality * force to use n_gpu_layers on simulator * add download buttons & expose llamaState.loadModel * update project.pbxproj * comment #Preview & fix editorconfig check * gitignore : xcode stuff * llama.swiftui : UX improvements * llama.swiftui : avoid data copy via "downloadTask" * llama.swiftui : remove model from project * llama : remove "mostly" from model infos * llama.swiftui : improve bench --------- Co-authored-by: jhen <developer@jhen.me>
2023-12-17	finetune : keep allocs alive until all allocations are done (#4486)	slaren

2023-12-17	server : disable llm logs if SERVER_VERBOSE is off (#3792)	olexiyb

2023-12-17	server : fix grammar being ignored (#4494)	AdithyanI
	Fix bug in identifying the grammar.
2023-12-17	server : fix possible ambiguity in content type charset (#4501)	Alexey Parfenov

2023-12-17	server : allow requests larger than 8K (#4500)	mzcu

2023-12-15	server : add optional API Key Authentication example (#4441)	ShadovvBeast
	* Add API key authentication for enhanced server-client security * server : to snake_case --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2023-12-14	ggml : remove n_dims from ggml_tensor (#4469)	slaren
	ggml-ci
2023-12-14	ggml : add ggml_row_size() (fixes llama out of space) (#4461)	LostRuins
	* Fixes "Not enough space in the context's memory pool" encountered on certain models, which seems to be caused by some imprecision related to the automatic casting of floating point values * do not cast to size_t, instead just use doubles * ggml : add ggml_row_size(), deprecate ggml_type_sizef() * ggml : fix row size compute to avoid overflows * tests : fix sizey -> sizez --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2023-12-13	server : fix handling of characters that span multiple tokens when streaming ↵	shibe2
	(#4446)
2023-12-12	server : tweak default sampling parameters (#4367)	kalomaze
	* Set a more typical Top P setting as the default * Update temp max
2023-12-12	english : use `typos` to fix comments and logs (#4354)	Richard Kiss

2023-12-12	server : fix local model name in server (#4420)	Vladimir Zorin

2023-12-10	Update README.md (#4388)	Yueh-Po Peng
	Fix small typo.
2023-12-07	llama : per-layer KV cache + quantum K cache (#4309)	Georgi Gerganov
	* per-layer KV * remove unnecessary copies * less code duplication, offload k and v separately * llama : offload KV cache per-layer * llama : offload K shift tensors * llama : offload for rest of the model arches * llama : enable offload debug temporarily * llama : keep the KV related layers on the device * llama : remove mirrors, perform Device -> Host when partial offload * common : add command-line arg to disable KV cache offloading * llama : update session save/load * llama : support quantum K cache (#4312) * llama : support quantum K cache (wip) * metal : add F32 -> Q8_0 copy kernel * cuda : add F32 -> Q8_0 copy kernel ggml-ci * cuda : use mmv kernel for quantum cache ops * llama : pass KV cache type through API * llama : fix build ggml-ci * metal : add F32 -> Q4_0 copy kernel * metal : add F32 -> Q4_1 copy kernel * cuda : wip * cuda : add F32 -> Q4_0 and F32 -> Q4_1 copy kernels * llama-bench : support type_k/type_v * metal : use mm kernel only for quantum KV cache * cuda : add comment * llama : remove memory_f16 and kv_f16 flags --------- Co-authored-by: slaren <slarengh@gmail.com> * readme : add API change notice --------- Co-authored-by: slaren <slarengh@gmail.com>
2023-12-07	train : fix #4227 (double free in ↵	Hongyu Ouyang
	examples/train-text-from-scratch/train-text-from-scratch.cpp) (#4351) On commit b1108 (44c117f4) xaedes added ggml_allocr * alloc = NULL; ... (many lines in between) if (alloc) { ggml_allocr_free(alloc); } Which is correct, but it's easy to lose context after many lines in between. On commit b1287 (0e76a899) xaedes made a big change. From here on, alloc is freed eagerly. alloc = ggml_allocr_new(...) ... (short lines of code) ggml_allocr_free(alloc) This happens a few times, but alloc is never set to NULL, and many lines below, we still have if (alloc) { ggml_allocr_free(alloc); } which causes a double-free.
2023-12-06	server : recognize cache_prompt parameter in OAI API (#4347)	Georgi Gerganov

2023-12-06	speculative : support `--color` (#4343)	stduhpf
	* speculative: add some colors * minor : add braces --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2023-12-05	sampling : custom samplers order (#4285)	MaggotHATE
	* Samplers sequence order w parameter * Cleaned commented code * Fixed formatting * Rewrote with unordered_map * Revert and rewrite, too many problems and safeguards would be needed * Fixed code style * Code style fixes according to review * More readable samplers input string, fixed help * Style fix in sampler_queue * Formatting fixes * Fixing whitespaces
2023-12-04	simple : update error message for KV cache check (#4324)	Daniel Bevenius
	This commit updates the error message that is printed when the KV cache is not big enough to hold all the prompt and generated tokens. Specifically it removes the reference to n_parallel and replaces it with n_len. Signed-off-by: Daniel Bevenius <daniel.bevenius@gmail.com>
2023-12-04	swift : fix concatenation method to avoid invalid UTF8 stringfication (#4325)	Miwa / Ensan

2023-12-04	swift : fix prompt tokenization logic (#4321)	Miwa / Ensan

2023-12-03	server : fix OpenAI API `stop` field to be optional (#4299)	Ed Lee
	(cherry picked from commit Mozilla-Ocho/llamafile@e8c92bcb84ae3bcbf0d617b7ee6a5413bcbd58af)
2023-12-03	py : add grammar to oai like api (#4294)	Rickard Edén

2023-12-01	llama : support optional tensors (#4283)	Georgi Gerganov

2023-12-01	swift : fix token_to_piece implementation (#4278)	Miwa / Ensan
	* Fix token_to_piece implementation in Swift * Fix errors
2023-12-01	ggml : add ggml_soft_max_ext (#4256)	Georgi Gerganov
	* metal : implement soft_max_ext * cuda : implement soft_max_ext * ggml : implement soft_max_ext (CPU) * batched-bench : print threads ggml-ci * metal : simplify soft_max encoding ggml-ci * cuda : use 512 threads for soft_max instead of 32 * ggml : update soft max cpu * cuda : do warp-based block reduce * cuda : increase max block size to 1024 * cuda : fix warp reduction initialization of shared mem * metal : warp-based reduction for soft max kernel * metal : warp-based reduce for rms_norm * metal : simplify soft max kernel ggml-ci * alloc : fix build with debug
2023-12-01	server : add --log-disable to disable logging to file (#4260)	Ziad Ben Hadj-Alouane
	* * add --log-disable to disable logging to file in the server example * * typo fix
2023-12-01	server : add single-client multi-prompt support (#4232)	Ziad Ben Hadj-Alouane
	* * add multiprompt support * * cleanup * * more cleanup * * remove atomicity of id_gen, and change lock_guard to unique_lock on completion requests * * remove all references to mutex_multitasks * Update examples/server/server.cpp Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com> * Update examples/server/server.cpp Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com> * Update examples/server/server.cpp Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com> * Update examples/server/server.cpp Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com> * * change to set --------- Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com>
2023-11-30	llava : ShareGPT4V compatibility (vision encoder only loading) (#4172)	John
	* ShareGPT4 compatibility (vision encoder only loading) Load only a CLIP vision encoder (as supplied by ShareGPT finetunes) Corrects the argument parsing for --img_mean and --img_std (which were previously not parsed but attempted to access) Defines defaults for img_mean and img_std which are equal to the llava 1.5 CLIP encoder, so you do not have to provide them * Update convert-image-encoder-to-gguf.py
2023-11-30	main : pass LOG_TEE callback to llama.cpp log (#4033)	Andrew Godfrey
	* main : Call llama_log_set to use LOG_TEE * tabs to spaces
2023-11-30	batched.swift : update README.md (#4214)	Miwa / Ensan
	docs: update how to run
2023-11-30	py : fix oai proxy (#3972)	rhjdvsgsgks
	* fix oai proxy fix generation not stoped while bot stop talking in chat mode fix possible `slot_id` not exist response for cors (and pre flight) * oai proxy: workaround for some client (such as Chatbox) * use stop as separator to replace hardcoded `\n`