Running a local AI model usually requires painfully high VRAM, and not everyone has the latest monster GPU with the coveted yet still absurd 20-plus gigabytes of VRAM. Still, it’s much more worthwhile to have compression enabled, since it would otherwise fill the limiting 8GB of VRAM. Gemma 2 (9B)The most powerful oneGemma 2 9B remains one of the largest models you can realistically run on a GPU with 8GB of VRAM. Keep your expectations in checkWith only 8GB of VRAM, expect to be severely bottlenecked in the long run. It’s unfortunate, but despite its popularity, 8 gigabytes is just not enough for local AI, and even gaming these days.