How to fit Qwen3.8-27B into 16GB of VRAM and run it on a single RTX 3080 card: the best quantizations and Llama.cpp flags I've found TL;DR: UD-IQ3_XXS with KV cache quantization … »
How to fit Qwen 3.6 35B A3B into 16GB of VRAM, & run it with Llama.cpp on an RTX 3080 The belly hangs over the belt, but it fits … »