My local LLM setup got faster after I stopped obsessing over parameter count
… For example, Granite 4 and Liquid AI's LFM2 both use a hybrid Mamba-transformer setup, and IBM claims it cuts memory use by up to 70% compared to a pure transformer at the same size. …