拓冰建站拓冰建站
首页 / 资讯中心 / 正文

当今改进 CNN、Transformer 还有出路吗?

一句话回答:出路不在”缝合”,而在六条根上主线——线性化与记忆、理论统一、稀疏注意力、硬件-算法协同、卷积的根上复兴、混合架构。2025–2026 年的最新证据恰恰表明,根上创新不仅没有枯竭,反而正在以前所未有的密度发生。一、问题本身:为什么”缝合”正在失效先说结论:“把前沿算法搬到下游任务”这条路正在以肉眼可见的速度枯竭。 这不是修辞,而是有明确证据的。证据一:ConvNeXt 证明”注意力收益”其实多是训练策略收益2022 年 Facebook AI Research 的 ConvNeXt(Liu et al., A ConvNet for the 2020s, arXiv:2201.03545, 已被引 9241 次)是最具说服力的一击。它什么都没发明——没有新算子、没有新模块——只是把 ResNet 按照 Transformer 的现代训练配方(AdamW 优化器、Label Smoothing、DropPath、Mixup/CutMix、Cosine 学习率调度、Layer Scale)重新训练了一遍。结果?添加图片注释,不超过 140 字(可选)图 1:ConvNeXt 原文 Figure 1。左:ImageNet-1K 训练;右:ImageNet-22K 预训练。气泡面积代表 GFLOPs。ConvNeXt 以 ConvNet 架构在精度上匹配或超越 Swin Transformer 和 ViT,而设计更简洁。原文: “The outcome of this exploration is a family of pure ConvNet models dubbed ConvNeXt. Constructed entirely from standard ConvNet modules, ConvNeXts compete favorably with Transformers in terms of accuracy and scalability, achieving 87.8% ImageNet top-1 accuracy and outperforming Swin Transformers on COCO detection and ADE20K segmentation, while maintaining the simplicity and efficiency of standard ConvNets.“(ConvNeXt, Liu et al. 2022)原文: “We demonstrate that a standard ConvNet model can achieve the same level of scalability as hierarchical vision Transformers while being much simpler in design.”这意味着什么?ViT(Dosovitskiy et al., 2020, 已被引 69024 次)和 Swin Transformer(Liu et al., 2021, 已被引 35253 次)当初宣称的”注意力机制本质优势”,其中很大一部分,其实是训练配方被误记到了注意力头上。ConvNeXt 之后,”纯 Transformer 优于纯 ConvNet”这个论断就再也不能不加限定地使用了。证据二:MetaFormer 证明算子本身不是根,宏观结构才是根同一年(2021 年 11 月),Yu et al. 提出了 MetaFormer(MetaFormer is Actually What You Need for Vision, arXiv:2111.11418, 已被引 1422 次)。他们的核心论点是:Transformer 的成功不应归功于注意力机制本身,而应归功于”宏观架构”——即”特征嵌入 → [Token Mixer + Token Processor] × L → 线性投影”这个框架。添加图片注释,不超过 140 字(可选)图 2:MetaFormer 原文 Figure 2。(a) PoolFormer 总体框架;(b) PoolFormer 块架构——它用 pooling 替代 attention 做 token mixing。最激进的一步是 PoolFormer:他们把 token mixer 从 attention 直接换成了 pooling(池化操作)——没有可学习参数,没有矩阵乘法——结果呢?原文: “We argue that the structure of the Transformer, i.e., the MetaFormer, is more important than the particular choice of token mixing operation.”PoolFormer-M48 在 ImageNet-1K 上达到了 82.5% 的 Top-1 准确率(见该论文 Table 2),而 Swin-Transformer-B(同样规模、相同数据)约为 81.0%–81.9%。也就是说:把 attention 换成 pooling,精度不仅没有下降,反而略有提升。这个发现的分量有多重?它意味着”发明一个更好的 token mixer”这条研究路线的价值,远没有”改进宏观架构”这条路线大。而缝合式论文恰恰是在前者这条路上反复撞墙。添加图片注释,不超过 140 字(可选)图 2b:MetaFormer 原文 Figure 3。ImageNet-1K 验证准确率 vs MACs 散点图。PoolFormer(粉色区域)与 MLP-Mixer(绿色)、Swin-Transformer(蓝色)、ResNet(黄色)在计算量-精度 Pareto 前沿上具有可比性,进一步证明 token mixer 的类型并非决定性因素。证据三:MambaOut 证明”新算子搬运”必须回答任务本质需求2024 年 5 月,Yu et al. 发表了 MambaOut(MambaOut: Do We Really Need Mamba for Vision?, arXiv:2405.07992)。这篇论文做的事情很极端:他们把 Mamba 块的核心——选择性状态空间模型(Selective SSM)——直接移除,只保留 Gated CNN 部分,结果在 ImageNet 分类上不仅没有下降,反而全面超越了所有视觉 Mamba 模型。添加图片注释,不超过 140 字(可选)图 3:MambaOut 原文 Figure 1。(a) Gated CNN 块(MambaOut 所用)与 Mamba 块的架构对比——Gated CNN 块本质上是 Mamba 块去掉 SSM 核心;(b) ImageNet 分类精度 vs MACs 散点图。MambaOut 模型在同等计算量下全面超越 Vision Mamba 和 PlainMamba。原文: “Mamba, an architecture with RNN-like token mixer of state space model (SSM), was recently introduced to address the quadratic complexity of the attention mechanism and subsequently applied to vision tasks. Nevertheless, the performance of Mamba for vision is often underwhelming when compared with convolutional and attention-based models. In this paper, we delve into the essence of Mamba, and conceptually conclude that Mamba is ideally suited for tasks with long-sequence and autoregressive characteristics. For vision tasks, as image classification does not align with either characteristic, we hypothesize that Mamba is not necessary for this task… our MambaOut model surpasses
分享:

看完干货,该让你的企业上线了

免费需求沟通 · 48 小时内出具建站方案 · 河南本地可上门