Self-hosting your LLMs (AIs): CPU+RAM Offloading, Mixture-of-Experts and Benchmarks
French version is available on LinuxFr Introduction This article is a follow-up to my previous entries: Self-hosting your AIs: General Principles Self-hosting your AIs: Hardware and Inference Optimization In this entry, we’re going to talk about running LLMs with little to no VRAM. In other words, we’re going to try and shove a large round peg into a small square hole. Like me, you might have noticed something strange: there are plenty of YouTube videos, LinkedIn posts, etc., telling you that running LLMs on a museum piece (like an Nvidia GTX 1060 6GB with DDR4) is super simple. They explain that it’ll run lightning-fast, work perfectly, and so on. Apparently, you just need to find the right “magic parameters,” and poof! You’ll never need Claude Fable 5 again! 🎉 ...