RASTP: Representation-Aware Semantic Token Pruning for Generative Recommendation with Semantic Identifiers

RecSys / IR Semantic Retrieval and Personalization SIL
2025年11月21日
生成式推荐系统通常利用语义标识符(Semantic Identifiers,简称SIDs),将每个物品表示为编码语义信息的令牌序列。然而,使用多个SIDs来表示物品ID会显著增加输入序列的长度,而序列长度是决定计算复杂度和内存消耗的主要因素。尽管现有研究主要集中在优化注意力计算和KV缓存上,我们提出了RASTP(Representation-Aware Semantic Token Pruning,即表征感知的语义令牌剪枝方法),该方法直接对输入序列中信息量较低的令牌进行剪枝。具体而言,RASTP通过结合语义显著性(基于表征幅度衡量)和注意力中心性(由累积注意力权重得出)来评估令牌的重要性。由于RASTP能够动态剪除信息量较低或无关的语义令牌,因此在三个真实世界的亚马逊数据集上的实验表明,RASTP可将训练时间减少26.7%,同时保持甚至略微提升了推荐性能。相关代码已开源,地址为 https://github.com/Yuzt-zju/RASTP。
Generative recommendation systems typically leverage Semantic Identifiers (SIDs), which represent each item as a sequence of tokens that encode semantic information. However, representing item ID with multiple SIDs significantly increases input sequence length, which is a major determinant of computational complexity and memory consumption. While existing efforts primarily focus on optimizing attention computation and KV cache, we propose RASTP (Representation-Aware Semantic Token Pruning), which directly prunes less informative tokens in the input sequence. Specifically, RASTP evaluates token importance by combining semantic saliency, measured via representation magnitude, and attention centrality, derived from cumulative attention weights. Since RASTP dynamically prunes low-information or irrelevant semantic tokens, experiments on three real-world Amazon datasets show that RASTP reduces training time by 26.7\%, while maintaining or slightly improving recommendation performance. The code has been open-sourced at https://github.com/Yuzt-zju/RASTP.
许愿