New AI startup could shrink server-sized models for use on iPhones
…The company has already been able to compress the 54GB Qwen 3.6 model to just 4GB. And, importantly, the technology that it uses doesn't impact the model's performance. Apple…
Tracked topic
Qwen3 is an AI model family developed by Alibaba, released as a set of large language models for natural-language tasks.
…The company has already been able to compress the 54GB Qwen 3.6 model to just 4GB. And, importantly, the technology that it uses doesn't impact the model's performance. Apple…
…Lukas Hauzenberger , , , , , , , Abstract KVpop learns optimal key-value cache eviction by directly supervising keep-or-drop decisions using future-attention targets, achieving high performance with reduced memory usage. Generated by Qwen/Qwen2…
…Notably, Qwen 3.7-Max can autonomously execute long-horizon agentic tasks—sustaining continuous operation for up to 35 hours and managing over 1,000 tool calls without performance degradation. Deeply optimized…
…auf einer Radeon RX 9060 XT für Gemma 4, Qwen3.5 und Qwen3.6 sowie auf einer Nvidia GeForce RTX 3090 für Gemma 4, Qwen3.6 und Qwen3.8. Die Geschwindigkeitsgewinne sind…
Looking for V100 users to share your config and it's performance. GPU: Tesla V100 PCIE 32Gb Qwen3.6 27B Q4_K_M + Q8_0 MTP 128K context length Pi coding agent llama.cpp model preset: [*] spec-default = 1 ctx-size = 131072…
Qwen3.8-Max (2.4T) is another massive contribution to the open weight community. On benchmarks, it performs closely to Kimi K3 and DeepSeek V4 flash across all categories and is better at coding and software tasks. Qwen3…
QwQ was genuine next-gen performance usable on local hardware, but the massive required context (it's reasoning style was akin to "if I say every possible word, I'll notice the right one!") kinda made it unusable for age…
Hey locallama! We recently open sourced a Qwen3-TTS 1.7B implementation that achieves 10 requests per second (RPS) and sub-50 ms p95 time-to-first-audio (TTFA) while maintaining real-time playback on 1 x H100. This exten…
From the description: "Reasoning-Medical-27B is designed for universal advanced medical reasoning in professional medicine, medical genetics, college biology/medicine, and clinical knowledge. The model was fine-tuned on …
SageMaker AI now supports serverless model customization for Qwen3.6 Posted on: May 14, 2026 Amazon SageMaker AI now supports serverless model customization for Qwen3.6 27B parameter model using supervised fine…
…Qwen-family models dominate the top-performing SLM tier in this setting. Qwen3-8B-Q4_K_M achieved the strongest overall SLM quality, reaching 0.72 correctness and 0.83 faithfulness, approaching…
…On BEIR , KaLM- Reranker -V1 achieves state-of-the-art performance, on par with strong industrial models such as the Qwen3- Reranker series; on MIRACL , despite not being extensively trained on multilingual…
…Qwen3 and DeepSeek V3, potentially becoming a new type of high-end CPU for the AI Agent era,” the post enthuses. Alibaba claims the machine’s single-core general-purpose performance “exceeded…
…First up, the Qwen 3.8-Max (2.8 trillion parameters) is second only to Moonshot's Kimi K3 in terms of performance. More intriguing, however, is the Qwen 3.8 (27…
…With Qwen3-4B as the backbone, our framework achieves the strongest aggregate performance on our benchmarks, outperforming larger proprietary LLMs (e.g., GPT, Gemini) and fixed-environment training baselines. We further analyze…