c/localllama · by Lantier@jlai.lu · 1yQwen/QwQ-32B · Hugging Face https://huggingface.co/Qwen/QwQ-32B18 points · 5 comments · view on lemmy.world
5 Comments
Even_Adder@lemmy.dbzer0.com · 3 pts · 1y
morrowind@lemm.ee · 2 pts · 1y
insane, absolutely insane
Suoko@feddit.it · 1 pts · 1y
Why insane? For quality, speed, size? I find the coder 1.5b and 3b light and good
morrowind@lemm.ee · 3 pts · 1y
It matches R1 in the given benchmarks. R1 has 671B params (36 activated) while this only has 32
Lantier@jlai.lu · 2 pts · 1y
GGUF quants are already out: https://huggingface.co/bartowski/Qwen_QwQ-32B-GGUF
JamonBear@sh.itjust.works · 1 pts · 1y
Yay! let's try
ollama run hf.co/bartowski/Qwen_QwQ-32B-GGUF:Q4_K_M/set parameter num_ctx 32768