I finally ditched Ollama after using llama.cpp's WebUI, and I'm not going back anytime soon
… When it came to using the Qwen 3.5:9b model, llama.cpp's WebUI generated exactly 69 tokens per second while Ollama lagged behind at 65 t/s. Ollama took around 15 seconds longer, but it was surprising seeing llama.cpp be blazing fast when I already thought Ollama was pretty sprightly. …