Chubby♨️@kimmonismus
45AI 编辑部评分,满分 100

GLM-5.3 发布:后训练带来网络安全能力跃升

2026-08-14 16:28· 1小时前
AI 导读

智谱发布 GLM-5.3,与 GLM-5.2 共用同一 743B 基础模型,全部性能提升来自扩展后训练(更多可执行环境、更长任务、更强验证器与强化学习)。网络评估中 ExploitBench 从 24.4% 升至 54.4%,两小时完成 105 个 ExploitGym 任务(GLM-5.2 为 29 个),权重计划两周后开源。作者预计美国政府将扩大监管框架以涵盖开源模型。

GLM-5.3 shows how much capability may still be hiding inside today's largest base models and how relevant post-training really is.

It uses the same base model (!) as GLM-5.2. Zai says the entire (!) improvement came from scaling post-training: more executable environments, longer tasks, stronger verifiers and more reinforcement learning.

Remember: Pre-training gives a model knowledge and raw problem-solving capacity. Post-training teaches it how to use that capacity: plan, call tools, test solutions, recover from failure and complete work over long horizons.

In cyber evaluations, GLM-5.3 moved from 24.4% to 54.4% on ExploitBench and completed 105 ExploitGym tasks in two hours, up from 29 for GLM-5.2. .

Its weights are scheduled for release in two weeks. However, numerous other open weight models will be released in the coming weeks:

-DeepSeek v4 Pro -Qwen3.8 27b -LTX 2.5 -Nemotron-Lighting -DeepSeek harness (just released, but harness isntead of a model) -Muse-Glimmer-30B (just released)

to name a few.

The US has meanwhile created classified cyber benchmarks and a voluntary pre-release process for "covered frontier models." What this release shows me, first and foremost, is that open models are continuing to move closer and closer to Frontier. And therefore, I believe that the US government will now further expand the regulatory framework to include open models.

That's why I'm even more excited for the ChatGPT "Astra" release. Because this model is *also* receiving a new (and more extensive) pre-training component, and we're currently seeing how much additional capability is enabled through post-training.

That's why this release is so significant; it demonstrates just how many areas for improvement are possible.

Chubby♨️WHAT: Zai just launched GLM-5.3, and its biggest leap may be in cybersecurity. The 743B base model remains unchanged (!) from GLM-5.2. Zai says the gains come e...

来源:Chubby♨️ · x.com

GLM-5.3 发布:后训练带来网络安全能力跃升

Chubby♨️ · @kimmonismus · X·2026-08-14 16:28·1小时前
AI 导读

智谱发布 GLM-5.3,与 GLM-5.2 共用同一 743B 基础模型,全部性能提升来自扩展后训练(更多可执行环境、更长任务、更强验证器与强化学习)。网络评估中 ExploitBench 从 24.4% 升至 54.4%,两小时完成 105 个 ExploitGym 任务(GLM-5.2 为 29 个),权重计划两周后开源。作者预计美国政府将扩大监管框架以涵盖开源模型。

GLM-5.3 shows how much capability may still be hiding inside today's largest base models and how relevant post-training really is.

It uses the same base model (!) as GLM-5.2. Zai says the entire (!) improvement came from scaling post-training: more executable environments, longer tasks, stronger verifiers and more reinforcement learning.

Remember: Pre-training gives a model knowledge and raw problem-solving capacity. Post-training teaches it how to use that capacity: plan, call tools, test solutions, recover from failure and complete work over long horizons.

In cyber evaluations, GLM-5.3 moved from 24.4% to 54.4% on ExploitBench and completed 105 ExploitGym tasks in two hours, up from 29 for GLM-5.2. .

Its weights are scheduled for release in two weeks. However, numerous other open weight models will be released in the coming weeks:

-DeepSeek v4 Pro -Qwen3.8 27b -LTX 2.5 -Nemotron-Lighting -DeepSeek harness (just released, but harness isntead of a model) -Muse-Glimmer-30B (just released)

to name a few.

The US has meanwhile created classified cyber benchmarks and a voluntary pre-release process for "covered frontier models." What this release shows me, first and foremost, is that open models are continuing to move closer and closer to Frontier. And therefore, I believe that the US government will now further expand the regulatory framework to include open models.

That's why I'm even more excited for the ChatGPT "Astra" release. Because this model is *also* receiving a new (and more extensive) pre-training component, and we're currently seeing how much additional capability is enabled through post-training.

That's why this release is so significant; it demonstrates just how many areas for improvement are possible.

Chubby♨️WHAT: Zai just launched GLM-5.3, and its biggest leap may be in cybersecurity. The 743B base model remains unchanged (!) from GLM-5.2. Zai says the gains come e...

来源:Chubby♨️· x.com