A Modder’s RTX 4080 Was Enough To Play AAA Games, But Not For Running LLMs, So He Integrated NVIDIA’s Tesla V100 At A Throwaway Price To Run 27B AI Models
…Adding the Tesla V100 grants the modder 32GB of usable VRAM, sufficient for running significantly improved AI models like Qwen 3.6 at 32 tokens/second Before you ask, it’s impossible…