llama.cpp

History

slaren 16bc66d947 llama.cpp : split llama_context_params into model and context params (#3301 ) * llama.cpp : split llama_context_params into model and context params ggml-ci * fix metal build * fix freq_base/scale default to model value * llama-bench : keep the same model between tests when possible * move n_threads to llama_context_params, add n_threads_batch * fix mpi build * remove kv_size(), cuda scratch fixes * remove low-vram option * add n_threads_batch to system info, refactor to get_system_info() * add documentation about --threads-batch to the READMEs * llama-bench fix * main : fix rope freq/scale warning * llama.cpp : add llama_get_model common : add llama_tokenize from model * remove duplicated ctx/model functions ggml-ci * cuda : print total VRAM used		2023-09-28 22:42:38 +03:00
..
.gitignore	llama : support input embeddings directly (#1910 )	2023-06-28 18:53:37 +03:00
CMakeLists.txt	cmake : install targets (#2256 )	2023-07-19 10:01:11 +03:00
embd-input-lib.cpp	llama.cpp : split llama_context_params into model and context params (#3301 )	2023-09-28 22:42:38 +03:00
embd-input-test.cpp	llama.cpp : split llama_context_params into model and context params (#3301 )	2023-09-28 22:42:38 +03:00
embd-input.h	examples : add compiler version and target to build info (#2998 )	2023-09-15 16:59:49 -04:00
embd_input.py	chmod : make scripts executable (#2675 )	2023-08-23 17:29:09 +03:00
llava.py	chmod : make scripts executable (#2675 )	2023-08-23 17:29:09 +03:00
minigpt4.py	chmod : make scripts executable (#2675 )	2023-08-23 17:29:09 +03:00
panda_gpt.py	chmod : make scripts executable (#2675 )	2023-08-23 17:29:09 +03:00
README.md	examples : fixed path typos in embd-input (#2214 )	2023-07-14 21:40:05 +03:00

README.md

Examples for input embedding directly

Requirement

build libembdinput.so run the following comman in main dir (../../).

make

LLaVA example (llava.py)

Obtian LLaVA model (following https://github.com/haotian-liu/LLaVA/ , use https://huggingface.co/liuhaotian/LLaVA-13b-delta-v1-1/).
Convert it to ggml format.
llava_projection.pth is pytorch_model-00003-of-00003.bin.

import torch

bin_path = "../LLaVA-13b-delta-v1-1/pytorch_model-00003-of-00003.bin"
pth_path = "./examples/embd-input/llava_projection.pth"

dic = torch.load(bin_path)
used_key = ["model.mm_projector.weight","model.mm_projector.bias"]
torch.save({k: dic[k] for k in used_key}, pth_path)

Check the path of LLaVA model and llava_projection.pth in llava.py.

PandaGPT example (panda_gpt.py)

Obtian PandaGPT lora model from https://github.com/yxuansu/PandaGPT. Rename the file to adapter_model.bin. Use convert-lora-to-ggml.py to convert it to ggml format. The adapter_config.json is

{
  "peft_type": "LORA",
  "fan_in_fan_out": false,
  "bias": null,
  "modules_to_save": null,
  "r": 32,
  "lora_alpha": 32,
  "lora_dropout": 0.1,
  "target_modules": ["q_proj", "k_proj", "v_proj", "o_proj"]
}

Papare the vicuna v0 model.
Obtain the ImageBind model.
Clone the PandaGPT source.

git clone https://github.com/yxuansu/PandaGPT

Install the requirement of PandaGPT.
Check the path of PandaGPT source, ImageBind model, lora model and vicuna model in panda_gpt.py.

MiniGPT-4 example (minigpt4.py)

Obtain MiniGPT-4 model from https://github.com/Vision-CAIR/MiniGPT-4/ and put it in embd-input.
Clone the MiniGPT-4 source.

git clone https://github.com/Vision-CAIR/MiniGPT-4/

Install the requirement of PandaGPT.
Papare the vicuna v0 model.
Check the path of MiniGPT-4 source, MiniGPT-4 model and vicuna model in minigpt4.py.