英文原文
MiMo-V2.6 is “simply” the best (for now). Despite its simple architecture design it’s currently No.1 in the open-weight benchmarks (weighted average).
With “simple,” I mean a classic Grouped Query Attention (GQA) with Sliding Window Attention (SWA) at a tiny 128-token window size.
So, that underlines one of the points I’ve been trying to make in recent months: most of the progress still comes from the data and post-training recipe improvements. Fancy attention variants are just mostly efficiency tweaks.
What are some of the training data improvements and recipe improvements? The MiMo team shared a pretty detailed technical report. Lots to carefully digest there, but in short, there are a few things that stood out:
- An increase in agent tasks; also training across different harnesses (the average DeepSWE pass@1 accuracy on held-out harnesses improved from approximately 50% → 66%).
- Better reward signals: they replaced a simple correctness verifier with an agentic grader that looks at the execution traces as well.
- Large RL batches (1,568 prompts × 16 rollouts = 25,088 trajectories) and 2.7–3.7 billion training tokens per update (unclear, though, what the predecessor used).
中文翻译
MiMo-V2.6 “简简单单”就是目前最好的模型。尽管它的架构设计很简单,但按加权平均计算,它目前在开放权重基准中排名第一。
这里的“简单”,是指经典的分组查询注意力(GQA)搭配滑动窗口注意力(SWA),而且窗口大小只有 128 个 token。
这也印证了我近几个月一直在强调的一点:大部分进步仍然来自数据和后训练方案的改进。花哨的注意力变体大多只是效率优化。
训练数据和训练方案具体有哪些改进?MiMo 团队分享了一份相当详细的技术报告,其中还有很多内容值得仔细消化;简要来说,以下几点尤其突出:
- 增加了 Agent 任务,并在不同的 harness 上训练(在留出的 harness 上,DeepSWE 的平均 pass@1 准确率从约 50% 提升到 66%)。
- 更好的奖励信号:他们用一个同时检查执行轨迹的 Agent 评分器,取代了简单的正确性验证器。
- 大规模 RL 批次(1,568 个提示 × 16 次 rollout = 25,088 条轨迹),每次更新使用 27 亿至 37 亿个训练 token(不过尚不清楚上一代模型使用了多少)。