Local LLM Inference Optimization: The Complete Guide

https://carteakey.dev/blog/local-inference/local-llm-optimization/

Good overview IMHO.

44 points · 2 comments · view on lemmy.world

2 Comments

theunknownmuncher@lemmy.world · 3 pts · 67d (1 reply)

Enabling XMP took my machine from roughly one-third speed back to normal.

Huge red flag. XMP is not designed for use with memory-intense workloads like running LLMs.

ComradePenguin@lemmy.ml · 2 pts · 67d

I have a PC that has been crashing a lot recently, but only when running long RAM+CPU and RAM+GPU intensive tasks. It uses about 60%+ of ny RAM.

I ran a memory test for 9 passes, no errors. No errors with the CPU only, and no errors with GPU only.

So it can be the XMP that crashes it?