Age | Commit message (Collapse) | Author | |
---|---|---|---|
2024-04-09 | BERT tokenizer fixes (#6498) | Jared Van Bortel | |
Key changes: * BERT conversion: fix abuse of LlamaHfVocab, do not set BOS or EOS * Nomic Embed conversion: pad vocab instead of slicing embedding tensor * llama_tokenize: handle added special tokens like HF does | |||
2024-03-23 | lookup: complement data from context with general text statistics (#5479) | Johannes Gäßler | |
* lookup: evaluation tools, use corpus/previous gens * fixup! lookup: evaluation tools, use corpus/previous gens * fixup! lookup: evaluation tools, use corpus/previous gens * fixup! lookup: evaluation tools, use corpus/previous gens * fixup! lookup: evaluation tools, use corpus/previous gens |