Skip to content

Companies & models

Zhipu GLM (Z.ai) Latest news

Follow Zhipu’s GLM models: flagship open releases, coding and reasoning improvements, and commercial and ecosystem developments.

9 picksPast 30 days: 0 itemsTotal: 51 items

Updated

Key moments

Since 2019
  1. ModelGLM-130B bilingual model released as open source
  2. ModelChatGLM-6B released
  3. ModelChatGLM3 released
  4. ModelGLM-4 released
  5. CompanyAdded to the US Entity List
  6. CompanyRebrands internationally as Z.ai
  7. ModelGLM-4.5 released
  8. ModelGLM-4.6 released
  9. ModelGLM-5.3 released with open weights

Zhipu GLM (Z.ai) top picks

Aug 27ThuItems 1–9
  1. Unsloth AI70

    GLM-5.3-Flash can run locally with Unsloth GGUF quantization on 128GB RAM

    Unsloth says GLM-5.3-Flash can run locally, with a 3-bit GGUF version needing 128GB of RAM and the 1-bit version working on 102GB of RAM or VRAM. The guide's table lists memory needs from 100GB at 1-bit to 650GB at BF16, and reports that the 1-bit quant keeps 71% of top-1% accuracy while being 85% smaller than BF16.

    Why it matters: The guide gives concrete memory requirements for each quantization level, which helps readers judge whether the model fits their hardware.

Aug 26Wed
  1. LMSYS Org65

    Zhipu's GLM-5.3-Flash adds native vision with day-0 SGLang support

    Z.ai released GLM-5.3-Flash, a 320B-A18B model, with day-0 support in SGLang, after appearing earlier as ox-alpha. The post calls it the first native multimodal model in the GLM-5 series and says it outperforms GLM-5.2 at one-tenth the cost, with stable 1M-token long-context performance.

    Why it matters: The post reports GLM-5.3-Flash's native multimodal design, its efficiency claims, and day-0 SGLang support, which bear on running it in practice.

Aug 25Tue
  1. Z.ai Release Notes62

    Z.ai releases GLM-5.3-Flash with native visual capabilities and hybrid architecture

    Z.ai has released GLM-5.3-Flash, a model with native visual capabilities that observe interfaces, rendering results, and interaction feedback across code, browsers, and GUIs. It uses a hybrid linear and sparse attention architecture with 320B total parameters and 18B activated, which the company says significantly reduces compute and KV-cache requirements. The release notes also describe support for office document and financial research workflows.

    Why it matters: The release notes give GLM-5.3-Flash's architecture, parameter counts, and cybersecurity findings, which make the model's scope concrete for comparison with earlier GLM releases.

  2. Z.ai (GLM) · new models on Hugging Face72

    Z.ai releases GLM-5.3 open weights with gains from post-training

    Z.ai released GLM-5.3 on Hugging Face, built on the same base model as GLM-5.2, with all gains coming from post-training. The source reports a 50% improvement over GLM-5.2 on Z.ai Code Bench and open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam, with a benchmark table comparing it against Kimi K3, DeepSeek-V4 Pro-0813, Qwen3.8-Max, and others.

    Why it matters: The source gives benchmark tables against GLM-5.2 and rival models, showing where the post-training gains concentrate in coding and cyber tasks.

Aug 17Mon
  1. Z.ai Release Notes63

    Z.ai releases GLM-5.3 with stronger coding and vulnerability discovery

    Z.ai's release notes announce GLM-5.3, which the company says delivers a 50% gain over GLM-5.2 on Z.ai Code Bench and reaches open-source SOTA on public benchmarks including Terminal Bench 3.0. The company also reports that GLM-5.3 matches Mythos 5 in white-box code review and vulnerability discovery, identifying 2,436 vulnerabilities in real-world targets, 1,097 of them medium- or high-severity. A separate GLM-5.3-Flash entry describes native visual capabilities and a hybrid architecture with 320B total and 18B activated parameters.

    Why it matters: The release notes show GLM-5.3's coding and cybersecurity gains, with a vulnerability count, letting readers compare it against Z.ai's prior GLM-5.x line and other coding models.

Jun 16Tue
  1. Z.ai (GLM) · new models on Hugging Face72

    Z.ai releases GLM-5.2 with 1M-token context and MIT open-source license

    Z.ai has released GLM-5.2, its flagship model for long-horizon tasks, which it says substantially improves on GLM-5.1 and supports a 1M-token context. The model adds IndexShare, which cuts per-token FLOPs by 2.9× at 1M context, and is released under the MIT open-source license.

    Why it matters: The source gives benchmark tables against named rival models and deployment settings, useful for judging where GLM-5.2 sits among current flagship models.

Jun 15Mon
  1. Z.ai Release Notes62

    Z.ai Release Notes: GLM-5.2 Adds 1M Lossless Context for Long Tasks

    Z.ai's release notes list GLM-5.2 as supporting 1M lossless context, with improved long-horizon task performance and reduced context drift and goal forgetting. The company says GLM-5.2 achieves open-source SOTA performance on coding and long-horizon task benchmarks. The page also includes the newer GLM-5.3 and GLM-5.3-Flash entries, which are listed above GLM-5.2.

    Why it matters: The page lists a dated series of Z.ai model releases, showing how the coding and long-horizon agent line has evolved from GLM-4.5 through GLM-5.2.

Apr 3Fri
  1. Z.ai (GLM) · new models on Hugging Face73

    Z.ai releases GLM-5.1, a flagship model for agentic engineering

    Z.ai has released GLM-5.1, its next-generation flagship model for agentic engineering, with stronger coding than GLM-5. The model is described as staying effective over longer agentic tasks, sustaining optimization over hundreds of rounds and thousands of tool calls. The release lists benchmark results including SWE-Bench Pro at 58.4 and Terminal-Bench 2.0 at 63.5, and local deployment is supported through SGLang, vLLM, xLLM, Transformers, and KTransformers.

    Why it matters: The release gives benchmark tables against several rival models, letting readers compare GLM-5.1's coding and agentic results with GLM-5 and frontier systems.

Feb 10Tue
  1. Z.ai (GLM) · new models on Hugging Face72

    Z.ai releases GLM-5, a 744B-parameter open model for agentic engineering

    Z.ai launches GLM-5, scaling from 355B to 744B total parameters with 40B active and pre-training data from 23T to 28.5T tokens. The model integrates DeepSeek Sparse Attention to reduce deployment cost and reports strong results on reasoning, coding, and agentic benchmarks against GLM-4.7, DeepSeek-V3.2, Kimi K2.5, and several frontier models.

    Why it matters: The source gives concrete scale, data, and benchmark comparisons against named frontier models, showing where GLM-5 sits among open-source and proprietary systems.