COMPUTE-EFFICIENT SCALING OF LARGE SEQUENTIAL RECOMMENDATION MODELS UNDER DYNAMIC USER BEHAVIOR

Authors

  • AoWei Shen (Corresponding Author) University of Washington, Seattle, WA 98195, USA.

Keywords:

Sequential recommendation, Scaling laws, Model growth, Continual learning, Concept drift, Compute efficiency

Abstract

Transformer-based sequential recommenders are increasingly scaled up and retrained as interaction logs grow, yet user behavior drifts between retraining cycles, and repeated from-scratch training of large models consumes compute that rarely yields proportional accuracy. This paper studies how model scale, update strategy, and training compute interact under a streaming protocol in which a model trained on earlier blocks of a chronologically ordered log serves the next block. On MovieLens-1M and Amazon Grocery, widening a SASRec-style model from 6.6K to 397K non-embedding parameters raises training compute by 4.4-13.3x relative to width 32 without a consistent accuracy gain, and models that are not refreshed lose 36% and 82% of their NDCG@10, respectively. Building on these observations, we propose SPROUT, a compute-aware update framework that warm-starts each update from a recency-aware replay stream, expands capacity with an exactly function-preserving width-doubling operator for pre-LayerNorm Transformers, and accepts growth only when a paired validation test shows a gain larger than one standard error. Over four update cycles, SPROUT stays within 1.3% of the best periodic-retraining accuracy on MovieLens-1M with 44% less cumulative compute than the retrained small model and 93% less than the retrained large model, and it achieves the highest accuracy on Amazon Grocery, 5.0% above the best retraining baseline, with 47% less compute than retraining the large model. Its growth-free configuration offers the most economical operating point, reaching 99.1% and 102.1% of the retraining accuracy with 33% and 51% of its compute.

References

[1] Hidasi B, Karatzoglou A, Baltrunas L, et al. Session-based recommendations with recurrent neural networks. In: Proceedings of International Conference Learn. Represent. (ICLR), 2016.

[2] Tang J, Wang K. Personalized top-N sequential recommendation via convolutional sequence embedding. In: Proceedings of ACM International Conference Web Search Data Mining (WSDM), 2018: 565-573.

[3] Yuan F, Karatzoglou A, Arapakis I, et al. A simple convolutional generative network for next item recommendation. In: Proceedings of ACM International Conference Web Search Data Mining (WSDM), 2019: 582-590.

[4] Kang W-C, McAuley J. Self-attentive sequential recommendation. In: Proceedings of IEEE International Conference Data Mining (ICDM), 2018: 197-206.

[5] Sun F, Liu J, Wu J, et al. BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer. In: Proceedings of ACM International Conference Inf. Knowl. Manage. (CIKM), 2019: 1441-1450.

[6] Li J, Wang Y, McAuley J. Time interval aware self-attention for sequential recommendation. In: Proceedings of ACM International Conference Web Search Data Mining (WSDM), 2020: 322-330.

[7] Zhou K, Wang H, Zhao W X, et al. S3-Rec: Self-supervised learning for sequential recommendation with mutual information maximization. In: Proceedings of ACM International Conference Inf. Knowl. Manage. (CIKM), 2020.

[8] Xie X, Sun F, Liu Z, et al. Contrastive learning for sequential recommendation. In: Proceedings of IEEE International Conference Data Eng. (ICDE), 2022: 1259-1273.

[9] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. In: Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 2017.

[10] Kaplan J, McCandlish S, Henighan T, et al. Scaling laws for neural language models. arXiv:2001.08361, 2020.

[11] Hoffmann J, Borgeaud S, Mensch A, et al. Training compute-optimal large language models. arXiv:2203.15556, 2022.

[12] Ardalani N, Wu C-J, Chen Z, et al. Understanding scaling laws for recommendation models. arXiv:2208.08489, 2022.

[13] Shin K, Kwak H, Kim S Y, et al. Scaling law for recommendation models: Towards general-purpose user representations. arXiv:2111.11294, 2021.

[14] Koren Y. Collaborative filtering with temporal dynamics. In: Proceedings of ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), 2009.

[15] Meng Z, McCreadie R, Macdonald C, et al. Exploring data splitting strategies for the evaluation of recommendation models. In: Proceedings of ACM Conference Recommender Syst. (RecSys), 2020: 681-686.

[16] Zhang Y, Feng F, Wang C, et al. How to retrain recommender system? A sequential meta-learning method. In: Proceedings of International ACM SIGIR Conference Res. Develop. Inf. Retrieval (SIGIR), 2020: 1479-1488.

[17] Mi F, Lin X, Faltings B. ADER: Adaptively distilled exemplar replay towards continual learning for session-based recommendation. In: Proceedings of ACM Conference Recommender Syst. (RecSys), 2020: 408-413.

[18] Lazaridou A, Kuncoro A, Gribovskaya E, et al. Mind the gap: Assessing temporal generalization in neural language models. In: Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 2021: 29348-29363.

[19] Chen C, Yin Y, Shang L, et al. bert2BERT: Towards reusable pretrained language models. In: Proceedings of Annual Meeting of the Association for Computational Linguistics (ACL), 2022: 2134-2148.

[20] Chen T, Goodfellow I, Shlens J. Net2Net: Accelerating learning via knowledge transfer. In: Proceedings of International Conference Learn. Represent. (ICLR), 2016.

[21] Gong L, He D, Li Z, et al. Efficient training of BERT by progressively stacking. In: Proceedings of International Conference Machine Learning (ICML), PMLR 97, 2019: 2337-2346.

[22] Wang J, Yuan F, Chen J, et al. StackRec: Efficient training of very deep sequential recommender models by iterative stacking. In: Proceedings of International ACM SIGIR Conference Res. Develop. Inf. Retrieval (SIGIR), 2021.

[23] Li Z, Hoiem D. Learning without forgetting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018, 40(12): 2935-2947.

[24] Rolnick D, Ahuja A, Schwarz J, et al. Experience replay for continual learning. In: Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 2019: 350-360.

[25] Ash J T, Adams R P. On warm-starting neural network training. In: Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 2020, 33: 3884-3894.

[26] Jean S, Cho K, Memisevic R, et al. On using very large target vocabulary for neural machine translation. In: Proceedings of Annual Meeting of the Association for Computational Linguistics (ACL), 2015: 1-10.

[27] Hastie T, Tibshirani R, Friedman J. The Elements of Statistical Learning. 2nd ed. New York, NY, USA: Springer, 2009.

[28] Harper F M, Konstan J A. The MovieLens datasets: History and context. ACM Transactions on Interactive Intelligent Systems, 2015, 5(4): 1-19.

[29] He R, McAuley J. Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering. In: Proceedings of International Conference World Wide Web (WWW), 2016: 507-517.

[30] Kingma D P, Ba J. Adam: A method for stochastic optimization. In: Proceedings of International Conference Learn. Represent. (ICLR), 2015.

Downloads

Published

2023-09-26

Issue

Section

Research Article

DOI:

How to Cite

AoWei Shen. Compute-Efficient Scaling Of Large Sequential Recommendation Models Under Dynamic User Behavior. World Journal of Information Technology. 2023, 1(1): 62-71. DOI: https://doi.org/10.61784/wjit4126.