Factorization Memory：高效RNN架构突破长文本建模

cs.AI updates on arXiv.org 5小时前

Factorization Memory：高效RNN架构突破长文本建模

本文提出了一种名为Factorization Memory的高效RNN架构，在短文本建模任务上与Transformer模型性能相当，同时在长文本场景中表现出更强的泛化能力。该模型基于Mamba-2，通过稀疏化形式优化了计算和内存效率，实现了在保持高性能的同时减少计算和内存消耗。

arXiv:2511.00315v1 Announce Type: cross Abstract: We propose Factorization Memory, an efficient recurrent neural network (RNN) architecture that achieves performance comparable to Transformer models on short-context language modeling tasks while also demonstrating superior generalization in long-context scenarios. Our model builds upon Mamba-2, enabling Factorization Memory to exploit parallel computations during training while preserving constant computational and memory complexity during inference. To further optimize model efficiency and representational capacity, we develop a sparse formulation of Factorization Memory that updates only a subset of recurrent states at each step while preserving the strong performance of its dense counterpart. To our knowledge, this represents the first RNN architecture that successfully combines sparse memory activation with competitive performance across both short and long-context settings. This work provides a systematic empirical analysis of Factorization Memory in comparison to Transformer and Mamba-2 architectures.

Fish AI Reader

AI辅助创作，多种专业模板，深度分析，高质量内容生成。从观点提取到深度思考，FishAI为您提供全方位的创作支持。新版本引入自定义参数，让您的创作更加个性化和精准。

FishAI

鱼阅，AI 时代的下一个智能信息助手，助你摆脱信息焦虑

联系邮箱 441953276@qq.com

相关标签

Factorization Memory RNN 长文本建模 Mamba-2 性能优化

相关文章

Sparse Maximal Update Parameterization (SμPar): Optimizing Sparse Neural Networks for Superior Training Dynamics and Efficiency

Node.js 最佳实践：开发人员指南

Webassembly：网络应用程序的近原生性能

This AI Paper from Databricks and MIT Propose Perplexity-Based Data Pruning: Improving 3B Parameter Model Performance and Enhancing Language Models

是时候向谷歌字体说再见了：缓存性能 (2020)

用于连接处理的简单、高效和稳健的哈希表

利用 Zig 的分配器

This AI Research Discusses Achieving Efficient Large Language Models (LLMs) by Eliminating Matrix Multiplication for Scalable Performance

Rails 上的异步 Ruby

使用 SIMD 指令更快地扫描 HTMLChrome 浏览器版