Evolving how LLMs are measured for Android: the next era of Android Bench
…Along with this change to our evaluation, in this July release we’re adding 8 new models ( Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2…
Tracked topic
Software engineers who’ve pitted GLM 5.2 against their own workflows report a wide range of results. “The main thing that I realized with [GLM 5.2], was that it can do more long-horizon tasks,” said Hasan, whose company hosts GLM 5.2 on North American infrastructure. Earlier open-weights models, he said, often lost the thread after around five to fifteen back-and-forth exchanges. “This one, I noticed that I could be using it for hours, and it would still have a coherent train of thought.” David Nix, a principal software engineer at the Denver-based MetaRouter, puts LLMs to work at both his day
Why Some Coders Now Reach for GLM 5.2 Before Frontier AI Models…Along with this change to our evaluation, in this July release we’re adding 8 new models ( Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2…
…model to analyze the attack, but instead used GLM 5.2, a recent release from Chinese AI lab Z.ai. Hugging Face’s security team didn’t access GLM 5.2 through…
…Hugging Face then turned to GLM 5.2, an open-weight model from Beijing-based Z . ai (formerly Zhipu AI), which had no such restrictions. The company's own July 16 disclosure…
…The kinds of volume I wanted AI models to run through would have blown through my current subscriptions to AI models (I have a ChatGPT Plus and GLM Coding Lite plan, which…
Claude Opus 4.5 vs. GLM-5.2
GLM-5.2 vs. Claude Opus 4.8: Full Comparison
GLM-5.2 vs. Claude Opus: Same Code, Less Than Half the Cost
GLM-5.2 (max) matches Claude Opus 4.8 on Harvey LAB-AA benchmark
Not directly LOCALLlama related but I thought it was interesting since Mistral and Z.ai are competitors, and more surprisingly they are pricing it (GLM-5.2) even cheaper than their current flagship model Mistral Medium 3…
…On this subset, frontier LLMs , including GPT-5.5, DeepSeek-V4-Pro, and GLM-5.1, reach only 30.00--45.67\%, a substantial drop from BrowseComp , while Korean LLMs released through…
…Hermes lets me swap models at any time I could switch to GLM or DeepSeek with ease M3 is what I'm using now, but I've also got GLM 5.2…
…On the comprehensive Verilog design problems (CVDP) benchmark, ACE-RTL with Nemotron 3 Ultra achieves a 97.1% average pass rate across nine agentic RTL task categories, outperforming models like GLM 5…
…We evaluate on WebShop , PinchBench , and Claw-Eval with Kimi-K2.5 , GLM-5 , and GPT-5.2 . SkillAdaptor improves over no-skill and skill-adaptation baselines on all three suites, with…
…Der chinesische KI-Anbieter Z.ai hat mit GLM-5.2 ein Open-Weight-Modell veröffentlicht, das bei der Erkennung von Sicherheitslücken mit Anthropics Opus 4.8 mithalten kann . Das ergaben IDOR…
…Hugging Face has said it resorted to using the open-source model GLM 5.2 to analyze the cyberattack because other frontier models were refusing to help due to built-in safety…