I stopped running the biggest local LLM that could fit, and a 2B model handles 90% of what I need
…So if you're not doing those, you're basically paying for capability you're not using. By paying I mean figuratively: it shows up as slower tokens per second, heavier disk…