Paper page - MotionVLA: Vision-Language-Action Model for Humanoid Motion
…However, many existing methods tokenize motion with a single shared codebook, forcing heterogeneous motion signals into the same quantization space. Our frequency-domain analysis of human motion data reveals a clear mismatch…