I tested a local LLM against a frontier cloud model, and the gap was smaller than I expected
… In one test, with 90,000 tokens of context and zero contamination, the local model gave the better answer. This isn't local models beating the cloud, because they don't overall . It's messier than that, but Qwen gives it some fantastic competition. …