ntn888

u/ntn888@lemmy.ml
12 posts · 57 comments

Recent posts

Recent comments

you can give "unloading some layers to RAM" a try though.. that way you can get your hands on the "usable" 31b models. browse around to find some good 31b ones.. GL

I didn;t try any 7b ones lately, they may be better fit for 16gb I think. I was able to try the 2b ones as I mentioned (on cpu). they are subpar. like mentioned the usable ones were 31b, I think you need atleast 24gb vram for most models though. maybe someone else can suggest better.

hey, thanks for your response.. yeah that's what I meant, the 2b models aren't usable in today's state, but more practical for everyday use if they work out..

I actually meant the 31b models are useful for my purpose. I don't do full-on agentic coding, just interactive chat/prompting. Example, I make good use for making linux shell scripts (as I don't know howto myself). Currently I use qwen3.5-flash via cloud. It's as good as the frontier models back then if not better..

on better Mastodon feeds · c/fediverse · 2 pts · 151d

hmm I see, I think I'll just try sticking. It must be working, the environment is much nicer than the other social media.. but the thought that I have to read all those posts..

on better Mastodon feeds · c/fediverse · 3 pts · 153d

"more" is not the keyword here.. "condensed" is what I'm looking for. I dont think you understand my point