The latest Gemma 4 models use a training trick to slash their on-device memory footprint
…unquantized QAT checkpoints, GPT-Generated Unified Format (GGUF), mobile-optimized, and Compressed Tensors. These models preserve “similar quality to bfloat16 while dramatically reducing the memory requirements to load the model,” according to…