I tested Google's new Gemma 4 12B on my 8GB GPU, and now I don't want to go back to smaller models
…So I cut the GPU offload to 32 and dropped context to 4K, which got the estimate to around 7.3 GB. Still failed. I updated LM Studio at this point because…
