you can give "unloading some layers to RAM" a try though.. that way you can get your hands on the "usable" 31b models. browse around to find some good 31b ones.. GL
I didn;t try any 7b ones lately, they may be better fit for 16gb I think.
I was able to try the 2b ones as I mentioned (on cpu). they are subpar. like mentioned the usable ones were 31b, I think you need atleast 24gb vram for most models though.
maybe someone else can suggest better.
funny I tried the 8B bonsai https://huggingface.co/prism-ml/Bonsai-8B-gguf
when loaded it takes ~7GB RAM!! When prompting it stalls my llama.cpp container (I'm running on a weak 4th gen i5)
hey, thanks for your response.. yeah that's what I meant, the 2b models aren't usable in today's state, but more practical for everyday use if they work out..
I actually meant the 31b models are useful for my purpose. I don't do full-on agentic coding, just interactive chat/prompting. Example, I make good use for making linux shell scripts (as I don't know howto myself). Currently I use qwen3.5-flash via cloud. It's as good as the frontier models back then if not better..
hmm I see, I think I'll just try sticking. It must be working, the environment is much nicer than the other social media..
but the thought that I have to read all those posts..
cool what was your hardware, and which qwen size you used? thanks
you can give "unloading some layers to RAM" a try though.. that way you can get your hands on the "usable" 31b models. browse around to find some good 31b ones.. GL
I didn;t try any 7b ones lately, they may be better fit for 16gb I think. I was able to try the 2b ones as I mentioned (on cpu). they are subpar. like mentioned the usable ones were 31b, I think you need atleast 24gb vram for most models though. maybe someone else can suggest better.
funny I tried the 8B bonsai https://huggingface.co/prism-ml/Bonsai-8B-gguf when loaded it takes ~7GB RAM!! When prompting it stalls my llama.cpp container (I'm running on a weak 4th gen i5)
Interesting thanks!
hey, thanks for your response.. yeah that's what I meant, the 2b models aren't usable in today's state, but more practical for everyday use if they work out..
I actually meant the 31b models are useful for my purpose. I don't do full-on agentic coding, just interactive chat/prompting. Example, I make good use for making linux shell scripts (as I don't know howto myself). Currently I use qwen3.5-flash via cloud. It's as good as the frontier models back then if not better..
ofcourse he's not blocking the blockade... he's merely imitating it :D
just wait till it gets federation.. it'll be the nail in the coffin for github!
hmm I see, I think I'll just try sticking. It must be working, the environment is much nicer than the other social media.. but the thought that I have to read all those posts..
"more" is not the keyword here.. "condensed" is what I'm looking for. I dont think you understand my point
Okay I'll give it a look thanks
it doesn't look like it's federated.. let alone linked to mastodon?
I searched for nodebb.. it is just forum hosting?
doesn't that make the situation worse with a busier feed??
is it global trending or personal feed? thanks
I wasn't even thinking about the interoperable fact, but indeed that is a bonus..
hmm yeah thanks. the interface is also rather inviting.. :)
yeah that comparison makes sense
hey! that thing is for domestic chinese??
yeah sure, but it's gotta be a collective thing.. btw what makes you stick to reddit over lemmy?