The biggest local LLM on your machine is useless if it can't call a single tool, no matter how many parameters it has
… Qwen's own published BFCL V4 results back this up: Qwen3.5 27B scores 68.5% and Qwen3.5 9B hits 66.1% and then there's a big drop: Qwen 3.5 4B drops to 50.3%, and Qwen 3.5 2B to 43.6%. …