Smaller, faster, safer: running Kimi and GLM at scale
… The effect is largest at low concurrency, where per-request latency matters most: Concurrent requests GLM FP8 tok/s GLM INT4 tok/s INT4 gain 1 60 92 +55% 8 425 513 +21% 16 683 825 +21% 32 994 1,267 +27% 64 1,672 1,933 +16% Prefill behaves differently. …