Open R1: Update #2
…Thanks! If you feel interested, can be reached out by chenyang.zhao@sglang.ai Would also be nice to include the performance of Qwen2.5-Math-Instruct and Qwen2.5-7B-Instruct…
Tracked topic
Qwen3 is an AI model family developed by Alibaba, released as a set of large language models for natural-language tasks.
…Thanks! If you feel interested, can be reached out by chenyang.zhao@sglang.ai Would also be nice to include the performance of Qwen2.5-Math-Instruct and Qwen2.5-7B-Instruct…
…startup PrismML has released Bonsai 27B, an extremely compressed AI model based on Alibaba’s Qwen3.6. The smallest version takes up just 3.9 GB, making it the first to fit…
…To demonstrate training utility, we collect rejection-sampled Qwen3.5 trajectories on the synthesized tasks and use them for supervised fine-tuning . Fine-tuning on these trajectories improves Qwen3.5-27B and…
…efficient deployment from edge devices to high-performance GPUs. This new generation of compact models supports a range of tasks, including: Reasoning: Strong performance on complex problem-solving tasks. Coding: Code generation…
Looking for V100 users to share your config and it's performance. GPU: Tesla V100 PCIE 32Gb Qwen3.6 27B Q4_K_M + Q8_0 MTP 128K context length Pi coding agent llama.cpp model preset: [*] spec-default = 1 ctx-size = 131072…
Qwen3.8-Max (2.4T) is another massive contribution to the open weight community. On benchmarks, it performs closely to Kimi K3 and DeepSeek V4 flash across all categories and is better at coding and software tasks. Qwen3…
We're releasing a fully quantized NVFP4 version of Qwen3.8-27B. The checkpoint was trained using quantization-aware distillation (QAD) with QUASAR, our new QAT algorithm. We used the original BF16 model as the teacher an…
QwQ was genuine next-gen performance usable on local hardware, but the massive required context (it's reasoning style was akin to "if I say every possible word, I'll notice the right one!") kinda made it unusable for age…
…Generated by Qwen/Qwen2.5-Coder-32B-Instruct Computer-use agents (CUAs) rely on visual observations of graphical user interfaces, where each screenshot is encoded into a large number of visual tokens…
…Qwen3-Coder-Next (80B MoE, 3B active) Qwen3.5-122B-A10B (122B MoE, 10B active) Devstral 2 123B (123B dense model) gpt-oss-120b (117B MoE, 5.1B active) Omnicoder-9B (9B…
…Article By Isabelle Liu Contributors Greg Gibby Related Blogs View All Blogs Run Qwen 3.8 27B on AMD Ryzen™ AI Max Agentic PCs and Radeon ™ GPUs Qwen3.8 27B arrives with…
…Udbhav Bamba , , , , Abstract DOT-MoE formulates dense layer decomposition as a differentiable optimal transport problem, enabling efficient training of sparse MoE models with improved performance retention. Generated by Qwen/Qwen2.5-Coder…
…After SFT on self-distilled data, the 3B model reaches performance comparable to, and in aggregate slightly above, Qwen3-Omni-30B-A3B-Instruct without using a stronger omni-modal teacher. These results…
…model pretrained from scratch demonstrates competitive performance on knowledge and reasoning tasks while highlighting differences in commonsense reasoning compared to autoregressive models. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Diffusion models…