Your old GPU can still run big LLMs – you just need the right tweaks
… Considering that quantization involves reducing the precision of model weights, Q8 can offer close to full accuracy of the model, but the VRAM hit is so massive that it makes sense to go for a higher parameter LLM and with heavier quantization. …