One Year Since the “DeepSeek Moment”
… Chinese Translation of this Article: https://huggingface.co/blog/vansin/one-year-since-the-deepseek-moment-cn great job @ irenesolaiman 👏 Great read! …
Tracked topic
… Chinese Translation of this Article: https://huggingface.co/blog/vansin/one-year-since-the-deepseek-moment-cn great job @ irenesolaiman 👏 Great read! …
… It’s inspiring to see how the DeepSeek R1 release helped expand the global open-source AI ecosystem, lowering barriers and sparking new innovation over the past year. …
… I have been trying to deploy deepseek-ai/DeepSeek-R1-Distill-Qwen-32B on inferentia with a context window higher than 4096 let's say MAX TOTAL TOKENS=8192 , but it seems there is no pre-compiled model for that. …
… MLA is not properly described in their paper, so it would be important to have code for this. · The code for the models are inside the model repositories, e.g. for V3: https://huggingface.co/deepseek-ai/DeepSeek-V3/blob/main/modeling deepseek.py Is it possible to contribute to this project? · Yes, … …
… View arXiv page View PDF GitHub 31 Add to collection Community Model: https://www.modelscope.cn/models/SLAIAITP/DeepSeek-V4-Flash-OR Code: https://github.com/SLAI-AITP/SLAI-T-Rex This is an automated message from the Librarian Bot . …
… Generated by Qwen/Qwen2.5-Coder-32B-Instruct Recently, end-to-end OCR models, exemplified by DeepSeek OCR, have once again thrust OCR into the spotlight. …
… On this subset, frontier LLMs, including GPT-5.5, DeepSeek-V4-Pro, and GLM-5.1, reach only 30.00--45.67%, a substantial drop from BrowseComp, while Korean LLMs released through Korea's Proprietary AI Foundation Model program obtain only 0.00--10.33%. …
… Headline results DeepSeek-R1-Distill-Qwen-1.5B, 4K budget, 5 math benchmarks : LEAD reaches 53.36 acc / 3714 tokens / +0.68 AES, the only method that improves accuracy over base while reducing length. 📄 Paper: https://arxiv.org/abs/2605.09806 💻 Code: https://github.com/CrazyMint/LEAD 🤗 Model: https… …
… To remain expressive, the indexer uses many query heads for example, 64 on DeepSeek-V3.2 that share the same selected token set; this multi-head design is precisely what makes the indexer the dominant cost on long contexts. …
… Don't forget, DeepSeek OCR also supports grounding OCR! wondering why minerU 2.5 model was not included in the comparison? …