Sebastian Raschka 这条帖子从架构角度拆解了刚公开身份的 GLM-5.3-Flash:此前以 Ox Alpha 名义出现的模型,采用稀疏 MoE、线性注意力与稀疏注意力混合的设计,并加入多流 mHC 残差路径和原生视觉编码器。它的价值不在于给出新的跑分结论,而在于把几项容易混淆的架构组件放到同一张图景里。

英文原文

Now we know: The popular Ox Alpha LLM was GLM-5.3-Flash…

Compared to GLM-5.2, this new GLM-5.3-Flash model uses:

  • a Kimi Linear-style 3:1 (super*) hybrid attention pattern with 34 Kimi Delta Attention layers (KDA) and 11 Multi-heat Latent Attention (MLA) / DeepSeek Sparse Attention (DSA) layers;
  • a scaled-down GLM-5.2-style sparse MoE backbone, going from 744B-A40B to 320B-A18B;
  • a DeepSeek V4-style mHC residual path with four parallel streams;
  • plus a native vision encoder (not shown).
  • "Super hybrid" because both KDA and MLA/DSA are "efficient" components. E.g., Kimi only uses KDA + full attention GQA, DeepSeek V3.2 uses DSA + full attention MLA.

PS: Sry for the excessive tech jargon. Explainers on all these components (MLA, DSA, KDA, mhC, etc.) in my LLM Architecture Gallery

PPS: Haha, maybe justification for getting that pricey Mac Studio M5 Ultra 256 GB / 512 GB to run this locally…

中文翻译

现在我们知道了:颇受关注的 Ox Alpha 大模型就是 GLM-5.3-Flash。

与 GLM-5.2 相比,新的 GLM-5.3-Flash 使用了:

  • 类似 Kimi Linear 的 3:1“超级”混合注意力模式,包括 34 个 Kimi Delta Attention(KDA)层,以及 11 个 Multi-head Latent Attention(MLA)/ DeepSeek Sparse Attention(DSA)层;
  • 缩小版的 GLM-5.2 风格稀疏 MoE 主干,从 744B 总参数、40B 激活参数缩减到 320B 总参数、18B 激活参数;
  • 类似 DeepSeek V4 的 mHC 残差路径,包含四条并行流;
  • 以及一个原生视觉编码器(图中未展示)。

之所以称作“超级混合”,是因为 KDA 和 MLA/DSA 都属于高效注意力组件。例如,Kimi 只使用 KDA 加完整注意力 GQA,而 DeepSeek V3.2 使用 DSA 加完整注意力 MLA。

附注:抱歉用了这么多技术术语。MLA、DSA、KDA、mHC 等组件的解释可以在作者的 LLM Architecture Gallery 中找到。

再附注:也许这能成为购买昂贵的 256 GB / 512 GB Mac Studio M5 Ultra、在本地运行模型的理由。

可核实事实与作者判断

Z.ai 的官方模型卡确认:GLM-5.3-Flash 是 GLM-5 系列首个原生多模态模型,采用 320B 总参数、18B 激活参数的稀疏 MoE;其总体设计结合稀疏注意力和线性注意力,并使用 Manifold-Constrained Hyper-Connections(mHC)。模型使用 30T token 多模态预训练语料,按 MIT 许可证发布。

Raschka 对 34 个 KDA 层、11 个 MLA/DSA 层、3:1 配比和四条并行残差流的表述,是他对模型架构的技术解读。官方模型卡支持混合注意力、MoE、mHC 和原生多模态这些总体事实,但页面正文没有逐项展开他使用的所有类比,因此不应把“Kimi 风格”或“DeepSeek V4 风格”理解为官方给出的等同结论。

如何验证与适用边界

官方模型卡列出 SGLang、vLLM、Transformers、KTransformers、Unsloth 和 TokenSpeed 等部署路径。实际测试时,应固定推理框架和版本、量化精度、上下文长度、reasoning_effort 以及 clear_thinking 参数,再记录显存或统一内存占用、首 token 延迟和解码吞吐。模型卡说明 reasoning_effort 支持 low、high、max,默认是 max;聊天场景建议显式设置 clear_thinking=true。

原帖本身没有提供可复现实验、运行命令、量化方案或 Mac Studio 实测数据。末尾关于 256 GB / 512 GB Mac Studio 的说法是作者的轻松评论,不能视为已经验证的硬件建议。官方公布的基准结果也应按其具体评测配置理解,不宜直接外推到本地量化部署。

官方模型卡:zai-org/GLM-5.3-Flash · Hugging Face
官方技术说明:https://z.ai/blog/glm-5.3-flash

原作者:Sebastian Raschka(@rasbt)
原帖:Sebastian Raschka on X: "Now we know: The popular Ox Alpha LLM was GLM-5.3-Flash... Compared to GLM-5.2, this new GLM-5.3-Flash model uses: - a Kimi Linear-style 3:1 (super*) hybrid attention pattern with 34 Kimi Delta Attention layers (KDA) and 11 Multi-heat Latent Attention (MLA) / DeepSeek Spa… / X