A SYSTEMS-AWARE FRAMEWORK FOR INTELLIGENT LARGE-SCALE RECOMMENDATION MODELS

Authors

  • AoWei Shen (Corresponding Author) University of Washington, Seattle, WA 98195, USA.
  • Shan Huang University of California San Diego, La Jolla, CA 92093, USA.

Keywords:

Recommender systems, Embedding compression, Memory hierarchy, Deep learning recommendation model, Systems-aware machine learning

Abstract

Embedding tables dominate the memory footprint of deep learning recommendation models, and their size already exceeds the capacity of the fast memory attached to accelerators. Existing compression methods reduce parameter counts without regard to where the parameters reside, while placement and caching systems move fixed tables between memory levels without changing their representation. This paper presents TierRec, a systems-aware framework that treats representation and placement as one decision. TierRec profiles the access skew of each categorical field together with an empirical fidelity profile of candidate representations, partitions the ID space into a full-dimensional hot tier, low-dimensional warm tiers with shared projections, and a compositional cold tier, and chooses the tier boundaries and dimensions by maximizing an access-weighted fidelity objective under a memory budget and, optionally, a slow-memory traffic constraint derived from a two-level memory model. Placement then follows access density so that compact rows occupy fast memory first. On MovieLens-20M with a DLRM-style ranking model, TierRec reaches a test AUC of 0.8199 at 8× compression and 0.8153 at 16× compression, matching or exceeding mixed-dimension embeddings and outperforming hashing, quotient-remainder and frequency-based double hashing by 0.0176 to 0.0560 AUC. The hierarchy-aware variant meets a slow-traffic bound of half the memory-only optimum at a cost of 0.0008 AUC. Trace-driven analyses on real and Criteo-scale synthetic workloads quantify the resulting gains in fast-memory hit ratio and access time.

References

[1] Rendle S. Factorization Machines. Proceedings of the IEEE International Conference on Data Mining (ICDM), 2010: 995-1000.

[2] Cheng HT, Koc L, Harmsen J, et al. Wide & Deep Learning for Recommender Systems. Proceedings of the 1st Workshop on Deep Learning for Recommender Systems (DLRS), 2016: 7-10.

[3] Guo H, Tang R, Ye Y, et al. DeepFM: A Factorization-Machine Based Neural Network for CTR Prediction. Proceedings of the 26th International Joint Conference on Artificial Intelligence (IJCAI), 2017: 1725-1731.

[4] Wang R, Fu B, Fu G, et al. Deep & Cross Network for Ad Click Predictions. Proceedings of ADKDD'17, 2017: 12.

[5] Zhou G, Zhu X, Song C, et al. Deep Interest Network for Click-Through Rate Prediction. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), 2018: 1059-1068.

[6] He X, Chua TS. Neural Factorization Machines for Sparse Predictive Analytics. Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2017: 355-364.

[7] Naumov M, Mudigere D, Shi H J M, et al. Deep Learning Recommendation Model for Personalization and Recommendation Systems. arXiv:1906.00091, 2019.

[8] Gupta U, Wu C J, Wang X, et al. The Architectural Implications of Facebook's DNN-Based Personalized Recommendation. Proceedings of the IEEE International Symposium on High Performance Computer Architecture (HPCA), 2020: 488-501.

[9] Acun B, Murphy M, Wang X, et al. Understanding Training Efficiency of Deep Learning Recommendation Models at Scale. Proceedings of the IEEE International Symposium on High Performance Computer Architecture (HPCA), 2021.

[10] Mudigere D, Hao Y, Huang J, et al. Software-Hardware Co-Design for Fast and Scalable Training of Deep Learning Recommendation Models. Proceedings of the 49th Annual International Symposium on Computer Architecture (ISCA), 2022.

[11] Zhao W, Xie D, Jia R, et al. Distributed Hierarchical GPU Parameter Server for Massive Scale Deep Learning Ads Systems. Proceedings of Machine Learning and Systems (MLSys), 2020.

[12] Adnan M, Maboud Y E, Mahajan D, et al. Accelerating Recommendation System Training by Leveraging Popular Choices. Proceedings of the VLDB Endowment, 2021, 15(1): 127-140.

[13] Sethi G, Acun B, Agarwal N, et al. RecShard: Statistical Feature-Based Memory Optimization for Industry-Scale Neural Recommendation. Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), 2022: 344-358.

[14] Weinberger K, Dasgupta A, Langford J, et al. Feature Hashing for Large Scale Multitask Learning. Proceedings of the 26th International Conference on Machine Learning (ICML), 2009: 1113-1120.

[15] Shi HJ M, Mudigere D, Naumov M, et al. Compositional Embeddings Using Complementary Partitions for Memory-Efficient Recommendation Systems. Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), 2020: 165-175.

[16] Ginart A A, Naumov M, Mudigere D, et al. Mixed Dimension Embeddings with Application to Memory-Efficient Recommendation Systems. Proceedings of the IEEE International Symposium on Information Theory (ISIT), 2021.

[17] Zhang C, Liu Y, Xie Y, et al. Model Size Reduction Using Frequency Based Double Hashing for Recommender Systems. Proceedings of the 14th ACM Conference on Recommender Systems (RecSys), 2020: 521-526.

[18] Yin C, Acun B, Liu X, et al. TT-Rec: Tensor Train Compression for Deep Learning Recommendation Model Embeddings. Proceedings of Machine Learning and Systems (MLSys), 2021.

[19] Kang WC, Cheng D Z, Yao T, et al. Learning to Embed Categorical Features Without Embedding Tables for Recommendation. Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2021.

[20] Liu S, Gao C, Chen Y, et al. Learnable Embedding Sizes for Recommender Systems. Proceedings of the International Conference on Learning Representations (ICLR), 2021.

[21] Zhao X, Liu H, Liu H, et al. AutoDim: Field-Aware Embedding Dimension Search in Recommender Systems. Proceedings of the Web Conference (WWW), 2021.

[22] Harper F M, Konstan J A. The MovieLens Datasets: History and Context. ACM Transactions on Interactive Intelligent Systems, 2015, 5(4): 19.

[23] Breslau L, Cao P, Fan L, et al. Web Caching and Zipf-Like Distributions: Evidence and Implications. Proceedings of IEEE INFOCOM, 1999: 126-134.

[24] Che H, Tung Y, Wang Z. Hierarchical Web Caching Systems: Modeling, Design and Experimental Results. IEEE Journal on Selected Areas in Communications, 2002, 20(7): 1305-1314.

[25] Duchi J, Hazan E, Singer Y. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization. Journal of Machine Learning Research, 2011, 12: 2121-2159.

[26] Kingma D P, Ba J. Adam: A Method for Stochastic Optimization. Proceedings of the International Conference on Learning Representations (ICLR), 2015.

Downloads

Published

2022-12-03

How to Cite

AoWei Shen, Shan Huang. A Systems-Aware Framework For Intelligent Large-Scale Recommendation Models. Journal of Computer Science and Electrical Engineering. 2022, 4(1): 36-44. DOI: https://doi.org/10.61784/jcsee4153.