Paper page - Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation
…and self-distillation reshapes the benchmark profile. After SFT on self-distilled data, the 3B model reaches performance comparable to, and in aggregate slightly above, Qwen3-Omni-30B-A3B-Instruct without using…