on Any way to prune LLMs? · c/localllama · 2 pts · 3yI don't know about that, but you could try GGML (llama.cpp). It has quantization up to 2-bits so that might be small enough.
I don't know about that, but you could try GGML (llama.cpp). It has quantization up to 2-bits so that might be small enough.