AFAIK there isn't a sparse 35B variant of 3.8, and I would not recommend running a dense 27B model on dual-channel DDR4 unless you are exceptionally patient. 3.6 is good enough for most things anyway, and I wouldn't expect OP to be doing heavy-duty coding tasks where you would notice the difference.
Any NPU/TPU you can buy is going to be essentially useless for either image diffusion or LLMs. The onboard RAM is both far too small and far too slow (LLM text generation speed relies on RAM speed first and foremost, and both LLMs and image models tend to be, you know, big), and USB isn't nearly fast enough to help with that, not to mention that software support is pretty much nonexistent. You'd be better off upgrading the GPU to a 3060ti 12gb or something.
P.S. A word of advice, consider using something other than Ollama. Llama.cpp in router mode or llama-swap support pretty much all of the functionality that Ollama does without being crap. Ik_llama.cpp is also nice if you have a CPU/Nvidia setup.
If you wanna make the most out of what you've got now, the LFM2.5 series of LLMs are quite good for the small size and fast inference speeds with sizes ranging from 0.2 billion to 8 billion parameters, though their low parameter count means that you'll probably wanna hook them up to some sort of web search or similar since they won't have a ton of general knowledge.
If you have at least 32GB of RAM, Qwen3.6 35B is quite a good general-purpose model that runs faster than its parameter count would suggest.
Eh, are you really using that many new models where sshing into your server / cding into your model folder, wgetting a url and configuring your interfaces every now and then is that much of an issue? I can't imagine using more than like one or two new models a month unless there's some insane string of releases or something.
Also, I mean, everyone's setup is different, but there's a significant amount of performance you're potentially leaving on the table by not using llama.cpp, potentially in the double-digit percentages. (Plus, if you have a fairly recent Nvidia setup and are willing to wait a bit for the latest models, ik_llama.cpp is a fantastic fork that I've found can get way better performance on most models than even llama.cpp.)
"Training" refers to a process where training data is passed backwards through a model in order to modify the model's weights. This happens before the model is deployed to the public. Search happens at (or right before) inference time (i.e. when the model is actually used) and does not modify the model weights.
Please don't use Ollama. Consider switching to llama.cpp instead, it's what Ollama uses under the hood and you'll have much finer control over what exactly it's doing so it'll be much easier to troubleshoot things like this.
"America's first car-free neighborhood" and "first car-free neighborhood," while both being strings present in the title, are not equivalent statements. I don't think it would be correct to omit "America's" in your comment.
Where in the article does it say it's the first car-free neighborhood ever? It just says it's the first in America, it even has a nod to how European cities tend to be built around walking in contrast to American cities.
Most messaging apps nowadays have read reciepts, a feature that shows the sender when the recipient opened the app and looked at their message. Being "left on read" means the recipient looked at the message but didn't bother to reply.
Happy to help! c:
AFAIK there isn't a sparse 35B variant of 3.8, and I would not recommend running a dense 27B model on dual-channel DDR4 unless you are exceptionally patient. 3.6 is good enough for most things anyway, and I wouldn't expect OP to be doing heavy-duty coding tasks where you would notice the difference.
Any NPU/TPU you can buy is going to be essentially useless for either image diffusion or LLMs. The onboard RAM is both far too small and far too slow (LLM text generation speed relies on RAM speed first and foremost, and both LLMs and image models tend to be, you know, big), and USB isn't nearly fast enough to help with that, not to mention that software support is pretty much nonexistent. You'd be better off upgrading the GPU to a 3060ti 12gb or something.
P.S. A word of advice, consider using something other than Ollama. Llama.cpp in router mode or llama-swap support pretty much all of the functionality that Ollama does without being crap. Ik_llama.cpp is also nice if you have a CPU/Nvidia setup.
If you wanna make the most out of what you've got now, the LFM2.5 series of LLMs are quite good for the small size and fast inference speeds with sizes ranging from 0.2 billion to 8 billion parameters, though their low parameter count means that you'll probably wanna hook them up to some sort of web search or similar since they won't have a ton of general knowledge.
If you have at least 32GB of RAM, Qwen3.6 35B is quite a good general-purpose model that runs faster than its parameter count would suggest.
Eh, are you really using that many new models where sshing into your server / cding into your model folder, wgetting a url and configuring your interfaces every now and then is that much of an issue? I can't imagine using more than like one or two new models a month unless there's some insane string of releases or something.
Also, I mean, everyone's setup is different, but there's a significant amount of performance you're potentially leaving on the table by not using llama.cpp, potentially in the double-digit percentages. (Plus, if you have a fairly recent Nvidia setup and are willing to wait a bit for the latest models, ik_llama.cpp is a fantastic fork that I've found can get way better performance on most models than even llama.cpp.)
Does llama.cpp router mode or llama-swap not meet your requirements? What doesn't work?
A word of advice, consider switching to something else.
Look, I hate Google as much as the next guy, but do you know what the phrase "picking your battles" actually means?
"Training" refers to a process where training data is passed backwards through a model in order to modify the model's weights. This happens before the model is deployed to the public. Search happens at (or right before) inference time (i.e. when the model is actually used) and does not modify the model weights.
Qwen2.5 (that model's base model) doesn't support image input
Indeed! It's all very vibes-based to be honest, though those two things do generally apply.
Please don't use Ollama. Consider switching to llama.cpp instead, it's what Ollama uses under the hood and you'll have much finer control over what exactly it's doing so it'll be much easier to troubleshoot things like this.
A trans woman (i.e. a person born as a man who later transitions to a woman) who likes women is homosexual, yes.
Homosexual = attraction to the same gender as one's own
Heterosexual = attraction to genders unlike one's own
https://sleepingrobots.com/dreams/stop-using-ollama/
Syntactically correct, but the way you phrased it implies that that's like a super duper niche usecase that no one uses when it really isn't.
OP is (presumably) vision impaired and accessibility still sucks on Linux.
Reminds me of a quote from Nevada by Imogen Binnie:
(Recited from memory, probably not word-for-word accurate.)
"America's first car-free neighborhood" and "first car-free neighborhood," while both being strings present in the title, are not equivalent statements. I don't think it would be correct to omit "America's" in your comment.
Where in the article does it say it's the first car-free neighborhood ever? It just says it's the first in America, it even has a nod to how European cities tend to be built around walking in contrast to American cities.
I mean, can you? By consuming her work or anything related to it, you're giving it clout which indirectly supports her.
Most messaging apps nowadays have read reciepts, a feature that shows the sender when the recipient opened the app and looked at their message. Being "left on read" means the recipient looked at the message but didn't bother to reply.