Z.AI has officially launched its latest flagship model, GLM-5.3, marking a significant breakthrough in coding capabilities and cybersecurity vulnerability discovery, while issuing a fresh challenge to AI leaders such as Anthropic and OpenAI.
According to an official statement from Z.AI, GLM-5.3 is built on the same 700-billion-parameter foundation model as its predecessor, GLM-5.2, with all performance gains stemming from expanded post-training phases. On Z.AI's proprietary Code Bench benchmark, GLM-5.3's coding ability improved by 50% compared to GLM-5.2. On the open-source coding benchmark Terminal-Bench 3.0, its score jumped from 4.6 to 28.3.
Z.AI described the model as being "born for coding" and ready to meet network defense challenges. GLM-5.3 also demonstrated an unexpected "emergent" leap in cybersecurity vulnerability mining, accumulating the identification of 2,436 security vulnerabilities across real-world codebases.
Z.AI stated that the model weights will be made public two weeks after the release, following the completion of safety assessments and reinforcement work. This launch coincides with DeepSeek's recent price hike for its V4 model, further strengthening Z.AI's competitive position in the domestic AI market.
Programming Capability Significantly Enhanced, Surpassing Claude Opus 4.8
The core breakthrough of GLM-5.3 lies in complex programming and long-chain tasks. Official data from Z.AI shows that on the highest difficulty setting of Z.AI Code Bench, GLM-5.3 completed 34.5% of tasks using approximately 75,000 output tokens, while GLM-5.2 achieved only a 23.4% completion rate despite consuming more tokens (96,000). On the high-difficulty setting, GLM-5.3 reached a 31.4% completion rate with about 50,000 tokens, surpassing Anthropic's Claude Opus 4.8 (29.5%, consuming 120,000 tokens). However, GLM-5.3 still lags behind Claude Fable 5, which achieved a 39.5% completion rate on the highest difficulty setting.
On public benchmarks, GLM-5.3 also performed strongly: its DeepSWE v1.1 score rose from 46.2 to 66.9, and its Agents' Last Exam score increased from 23.8 to 28.5. To achieve these advancements, Z.AI significantly expanded the scale and diversity of its training environment during the post-training phase, introducing tasks more closely resembling real-world engineering scenarios—some with workloads equivalent to several days of work for an experienced engineer. Z.AI also built an automated pipeline for end-to-end synthesis of training environments and reinforcement learning reward signals.
Cybersecurity Capabilities Emerge, Over Two Thousand Vulnerabilities Disclosed
In the field of cybersecurity, GLM-5.3 has demonstrated an unexpected leap in capability. After Z.AI introduced vulnerability discovery data and training environments, the model not only identified single defects but also began performing comprehensive reasoning across multiple vulnerability exploitation stages, forming complete exploit chain plans.
On the CyberGym benchmark, GLM-5.3 scored 84.5%, surpassing Anthropic's Mythos 5 (83.8%) and OpenAI's GPT-5.6 Sol (83.6%) to take the top spot. On ExploitBench, which requires deeper vulnerability reasoning, GLM-5.3's score jumped from 24.4% for GLM-5.2 to 54.4%, more than doubling; however, it still lags significantly behind Mythos 5 (78.0%) and GPT-5.6 Sol (76.5%). In the ExploitGym test, GLM-5.3 completed 105 exploitation tasks in two hours and 130 tasks in six hours, compared to GLM-5.2's 29 and 39 tasks, respectively; Mythos 5 led with 181 and 247 tasks.
Z.AI stated that these capabilities have extended to real-world applications. Since GLM-5.2, Z.AI has collaborated with multiple security teams in China to run model tests on real-world codebases. After expert review, screening, and deduplication, the model identified 2,436 vulnerabilities across 269 projects, including 1,097 medium-to-high-risk vulnerabilities. These spanned system kernels, operating systems, browser engines, open-source infrastructure, web applications, and network protocols. Some vulnerabilities had persisted in codebases for decades, with the earliest dating back approximately 40 years. Z.AI has established a "Z.AI Security Disclosure Ledger" for ongoing public disclosure.
Open-Source Strategy Coupled with Competitive Landscape Shifts Opens Market Opportunities
On the commercial front, the timing of GLM-5.3's release is strategically significant. DeepSeek announced on Thursday that it would raise prices for its V4 Flash and Pro models by up to four times. Although the overall pricing of V4 Pro remains lower than GLM-5.2, this move has objectively created a window for Z.AI to attract users and developers. Artificial Analysis assigned DeepSeek's top model the same intelligence score of 53 as GLM-5.2, with both trailing behind Kimi K3, Claude Fable 5, and GPT-5.6.
Z.AI plans to open the model weights of GLM-5.3 under a permissive license to attract a broader developer community. This strategy follows the same path as GLM-5.2, which saw Z.AI's market value surge to $137 billion, surpassing internet giants like Pinduoduo and NetEase, before retreating to around $80 billion—still a more than tenfold increase from its listing in Hong Kong in January of this year.
According to a previous Bloomberg report, Z.AI has completed construction of a data center equipped with at least 10,000 domestically produced chips for the development and inference of the GLM series. Its annual recurring revenue (ARR) reached $1 billion in July of this year.