Two old GPUs I salvaged are doing more AI work than a brand new $2000 card, and I won't be upgrading anytime soon
…But this limitation only applies to conventional LLMs, not Mixture-of-Experts models. Rather than relying on the entire model to process prompts like conventional LLMs, MoE clankers use a router + experts…
