Pith. sign in

REVIEW 5 major objections 5 minor 114 references

Lightweight Embeddings with Graph Rewiring for Collaborative Filtering

T0 review · 5 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A framework called LERG claims to make GNN-based recommenders edge-deployable by quantizing a compositional embedding codebook and pruning low-contribution graph nodes, cutting the embedding table to 1.76–17.12 MB while reporting the best…

desk verdict Solid engineering paper with a load-bearing rewiring premise that the ablation does not actually test. read the letter →

arxiv 2505.18999 v1 pith:7YCEAND7 submitted 2025-05-25 cs.IR

classification cs.IR
keywords lightweightrecommendersystemscompositionalembeddingcodebookquantizationgraphrewiringcollaborativefilteringedgedeploymentbinaryintegerprogrammingcompression
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that the two main bottlenecks to running graph-based recommender systems on small devices—embedding-table storage and graph-propagation computation—can be attacked with one framework. Its proposed LERG quantizes a compositional codebook of shared meta-embeddings so that more meta-embeddings fit in the same storage budget, then rewires the user–item interaction graph by pruning entities judged to contribute little to collaborative signal propagation. On Yelp2020, Amazon-book, and the industry-scale iFashion dataset, it reports the best accuracy among lightweight baselines while cutting the embedding table to 1.76–17.12 MB and lowering peak memory and MACs below all compared methods. The practical stake is that GNN-based recommenders could be fine-tuned and run directly on edge devices rather than querying a server.

What carries the argument

The central object is the quantized compositional embedding table: a dense codebook of $c$ meta-embeddings stored as $b$-bit integers with per-meta-embedding step sizes, plus a sparse assignment matrix $S$ that composes each entity's embedding from an anchor meta-embedding and an auxiliary meta-embedding. The second load-bearing mechanism is the rewired propagation graph $A'$: starting from pretrained graph-propagated embeddings $H^{\mathrm{pretrain}}$, the method builds the entity–entity similarity matrix $B$, relaxes a binary integer program that picks the top-$m$ entities by total similarity, and repairs zero rows by connecting isolated entities to indirect neighbors up to $T$ hops. Together these pieces keep recommendation accuracy while shrinking storage to $O(c(\frac{b}{32}d + 1)) + O(rd + (N-m))$ and cutting propagation MACs in proportion to the retention ratio.

What would settle it

Compare LERG against a variant that prunes the same number of entities by interaction degree instead of by the row-sum of $B$: if the degree-pruned variant matches or beats LERG on held-out NDCG or Recall on any of the three datasets, the claim that row-sum similarity captures propagation contribution fails. A cheaper check is to inspect the pruned set: if it contains many high-degree hubs or nodes that are the only bridge between dense communities, the contribution measure is misspecifying importance.

Watch

Extended reading notes

Core claim

In LERG, the full embedding table is replaced by a quantized compositional codebook $\bar{E}^{\mathrm{meta}} \in \{-2^{b-1},\dots,2^{b-1}-1\}^{c \times d}$ stored in low-bit integers together with a learnable step-size vector $\Delta$, combined with a fixed highly sparse assignment matrix $S$; the full table is recovered as $\hat{E} = S(\bar{E}^{\mathrm{meta}} \times \Delta)$. Quantization-aware pretraining on the full graph is followed by a graph-rewiring step: entities are scored by the row sums of the similarity matrix $B = H^{\mathrm{pretrain}} H^{\mathrm{pretrain}\top}$, a relaxed binary integer program selects the $m$ most impactful entities, and multi-hop rewiring fills in neighbors for any node left isolated. Only retained entities are fine-tuned on the rewired graph; pruned entities receive imputed embeddings from a small placeholder codebook built by clustering their pretrained embeddings. The paper reports that this pipeline outperforms prior lightweight methods, including its predecessor LEGCF, on all three datasets under comparable storage, while using the least peak memory and the fewest MACs.

Load-bearing premise

The load-bearing premise is that an entity's contribution to collaborative signal propagation is correctly measured by the total similarity between its pretrained embedding and all other entities, so pruning the lowest-scoring entities before fine-tuning is safe.

Editorial extensions

If this is right

  • GNN-based recommenders can be deployed on resource-constrained devices with embedding tables of a few megabytes instead of hundreds of megabytes, with modest accuracy loss relative to full-dimensionality settings.
  • On-device fine-tuning becomes feasible: because only the retained entities' embeddings are updated, peak memory and MACs are capped, allowing adaptation to new interactions without a server round-trip.
  • The retention ratio $m$ acts as a tunable dial between accuracy and computational cost, letting a single pretrained model serve different hardware budgets by regenerating the rewired graph and fine-tuning.
  • The framework transfers across base GNN recommenders, so the storage and propagation savings are not tied to one architecture.
  • Using INT16 rather than INT8 or INT4 for the quantized codebook matters most on large-scale datasets, indicating a precision-versus-expressiveness trade-off that can be set per deployment.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The row-sum centrality score behind Eq. 9 is only one possible importance measure; a message-passing-aware centrality such as Personalized PageRank might make the rewiring more robust on graphs where embedding-space similarity does not align with propagation influence.
  • The same rewiring recipe could be applied to other graph learning tasks, such as node classification on large graphs, because the pretrained propagated embeddings already encode collaborative semantics and the storage/computation gains are task-agnostic.
  • A testable extension is to make rewiring adaptive: recompute the BIP scores from fine-tuned embeddings as new interactions arrive on-device and rewire incrementally, rather than fixing the graph at deployment time.
  • The quantization-bit results suggest a per-dataset bit-selection rule could squeeze further storage savings without retraining, since INT8 suffices for smaller datasets while INT16 is needed for the industry-scale one.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper proposes LERG, a lightweight GNN-based recommender system that combines a quantized compositional codebook (meta-embeddings quantized to INT4/8/16 with learned step sizes) with graph rewiring. After pretraining the quantized codebook on the full interaction graph, the method forms an entity-entity similarity matrix from the pretrained propagated embeddings, solves a relaxed binary integer program to select a set of 'high-contribution' entities, prunes the adjacency-matrix columns of the remaining entities, and rewires zero-row entities to multi-hop neighbors. Pruned entities receive imputed embeddings from a small placeholder codebook obtained by K-means. The system is evaluated on Yelp2020, Amazon-book, and Alibaba-iFashion, reporting lower embedding storage and MAC counts than existing lightweight baselines while claiming better recommendation accuracy.

Significance. If the empirical claims hold, LERG is a practically useful contribution to on-device GNN recommendation: storage is reduced to 1.76-17.12 MB, peak memory and MAC reductions are consistent across datasets, and the method is validated at industry scale and across three base recommenders. The paper also contains end-to-end quantization-aware training, ablations of fine-tuning, rewiring, and placeholder imputation, and hyperparameter studies. However, the central rewiring mechanism is justified by an unvalidated proxy objective, the BIP formulation is mathematically mischaracterized, and all quantitative claims lack variance estimates; these issues must be resolved before the comparisons can be taken at face value.

major comments (5)
  1. [§3.6, Eq. (14)] The propagation equation in Eq. (14) multiplies an m×d input H_retain by the N×N rewired adjacency matrix A′ defined in Eq. (13) and Algorithm 1, while the degree matrix D is stated to be m×m. These dimensions are inconsistent and the matrix product is not defined as written. If propagation is performed only on the retained subgraph, A′ must be replaced by its m×m restriction, or the retained rows/columns must be explicitly extracted. Please define the exact matrix used in Eq. (14) and state how the symmetric normalization in Eq. (2) is adapted to the directed rewired graph.
  2. [§3.5, Eq. (10)] The optimization in Eq. (10) is a separable linear selection problem: its optimum is simply the set of m entities with the largest row sums R_j = Σ_k B_jk. Consequently, the claims that the problem is NP-complete, that an LP relaxation (Eq. 11) is necessary, and that the simplex algorithm supplies a nontrivial solution are incorrect; sorting the row sums suffices. More substantively, the objective does not measure propagation contribution in the rewired graph: R_j aggregates similarity to all entities including those that will be pruned, and the objective does not maximize similarity among the retained set. Please reformulate the objective toward retained-set similarity (e.g., Σ_{j,k∈N_retain} B_jk) or provide direct empirical evidence that the row-sum proxy tracks the actual effect of entity removal on recommendation accuracy.
  3. [Table 4, w/o BIP ablation] The only comparison for the rewiring criterion is against random pruning. At retention ratios of 0.5 and 0.1, any non-random rule is expected to beat random pruning, so this ablation cannot validate the specific row-sum premise in Eq. (10). Add comparisons with degree-based pruning, pruning by embedding norm, and an oracle rule that removes the entity whose deletion causes the largest validation-loss drop, and report these across the retention ratios used in Table 4.
  4. [Tables 3-6, Figures 2-4] All experiments appear to be single runs; no standard deviations, confidence intervals, or significance tests are reported. Several margins supporting the central claims are very small (e.g., Table 3: LERG vs LEGCF on Amazon-book N@10 is 0.0146 vs 0.0142; Table 4: default vs w/o Fine-tuning on iFashion at 0.7 is 0.0045 vs 0.0043). Please report averages over multiple seeds and perform significance testing, otherwise the claim of 'best performance across all three datasets' is not robustly supported.
  5. [§4.1.4 vs Fig. 3a] The default codebook sizes in §4.1.4 (c=2,000 for Yelp2020 and Amazon-book, c=10,000 for iFashion) contradict the hyperparameter study in Fig. 3a, which reports that c=5,000 yields optimal performance on all three datasets, and the text immediately adds that c=2,000 outperforms c=5,000 on iFashion. The main results in Table 3 are therefore not reported at the tuned optimum. Please reconcile the default with the sensitivity analysis or explicitly justify the discrepancy.
minor comments (5)
  1. [§4.3] The retention-ratio list is printed as {0.7, 0.5, 0,1}; it should read {0.7, 0.5, 0.1}.
  2. [§4.7] Several typos should be fixed: 'pertaining process' should be 'pretraining process', 'finefining' should be 'fine-tuning' in §4.2, and 'on-sever' should be 'on-server' in §3.6.
  3. [§4.1.2, Table 3] The baseline is named 'Post4bits' in the text and 'Post4Bits' in the table; please use one consistent name.
  4. [References] The manuscript still contains ACM placeholder text ('Conference acronym', 'September 2018', 'Received 20 February 2007'), which must be replaced before any publication.
  5. [Fig. 2] Figure 2 packs two plots per dataset panel but the caption does not describe the symbol and color encoding for the left and right subplots; please add a legend or detailed caption.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: LERG's performance claim is an empirical comparison; the quantization and graph-rewiring components are implemented from stated equations and tested against external baselines.

full rationale

The paper's central claim is empirical: LERG outperforms baselines on three datasets under matched storage budgets (Tab. 3). The method builds on the authors' prior LEGCF, but LEGCF is re-run as a baseline in the same experimental setup and the current work also ablates the assignment-initialization method (Fig. 3c), so the inheritance is not a load-bearing self-citation. The quantization storage reduction (Sec. 3.3) is arithmetic from Eqs. 6-7 and the stated bit width, not a fitted-then-predicted quantity. The graph-rewiring selection (Eq. 10) is a heuristic based on pretrained embeddings; it is not defined in terms of the reported accuracy metric, and the w/o BIP ablation compares it against random pruning, which is a substantive empirical test even if the premise is debatable. No target quantity is fitted to a subset and then reported as a prediction, and no uniqueness theorem is imported from the authors. Concerns that Eq. 10's row-sum objective does not measure propagation contribution are about validity and evidence strength, not circularity. Therefore the derivation chain is self-contained for the purpose of this review.

Assumptions & free parameters 7 free parameters · 5 assumptions · 2 invented entities

The method depends on several hand-chosen hyperparameters and domain assumptions. Most are standard in the lightweight recommendation and quantization literature, but the similarity-based pruning proxy is the most consequential assumption, since it decides which entities get removed from graph propagation.

free parameters (7)
  • codebook size c = default 2,000 (Yelp/Amazon), 10,000 (iFashion); Fig. 3a suggests 5,000
    Model capacity hyperparameter tuned per dataset; default and reported optimum disagree.
  • retention ratio = 0.7 default; 0.1 to 1.0 studied
    Controls the accuracy/efficiency trade-off; chosen by validation.
  • quantization bit width b = 16 default; 4 and 8 tested
    Sets storage per codebook element; performance varies by dataset.
  • placeholder codebook size r = 500 (Yelp/Amazon), 2,000 (iFashion)
    Number of K-means centroids for pruned entities; tuned.
  • rounding boundary o = 0.5
    Hand-chosen threshold for rounding the LP solution to binary variables.
  • anchor weight w* = 0.9
    Carried over from the LEGCF hyperparameter study.
  • max rewiring hops T = 4
    Chosen to recover all-zero rows; no sensitivity study is provided.
assumptions (5)
  • domain assumption Pretrained entity-entity similarity is a valid proxy for an entity's contribution to graph propagation (Eq. 9-10)
    The pruning decision, and thus the graph structure used in fine-tuning, rests on this heuristic. The w/o BIP ablation supports it empirically but only on three datasets.
  • domain assumption LSQ quantization-aware training with STE preserves enough fidelity to maintain recommendation accuracy
    Standard in the quantization literature; the paper relies on it for the codebook.
  • domain assumption METIS graph partitioning provides a good initialization of anchor meta-embeddings
    Taken from LEGCF; Fig. 3c shows graph-partition initialization outperforms random.
  • domain assumption K-means centroids of pretrained pruned embeddings adequately impute pruned entities
    Eq. 15-16; ablation without C_prune shows large drops at low retention ratios.
  • standard math The LP relaxation of the entity selection problem has the same optimum as the BIP
    For a single cardinality constraint with a linear objective this is true, but the paper does not state it; instead it frames BIP as NP-complete and uses rounding.
invented entities (2)
  • placeholder meta-embedding codebook C_prune
    purpose: Provides cheap imputed embeddings for pruned entities so their full embeddings need not be stored or propagated.
    A model component introduced by the paper; no external falsifiable prediction.
  • rewired propagation graph A'
    purpose: Sparsified directed graph used on edge devices to reduce MACs.
    Constructed by pruning columns of the interaction graph and adding indirect neighbors; its benefit is measured only through downstream accuracy.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Lightweight Embeddings with Graph Rewiring for Collaborative Filtering." pith.science (2026). https://pith.science/paper/7YCEAND7

@misc{pith2026250518999,
  author       = {Pith},
  title        = {Pith review of: Lightweight Embeddings with Graph Rewiring for Collaborative Filtering},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7YCEAND7}},
  note         = {Machine review of arXiv:2505.18999}
}
read the original abstract

As recommendation services scale rapidly and their deployment now commonly involves resource-constrained edge devices, GNN-based recommender systems face significant challenges, including high embedding storage costs and runtime latency from graph propagations. Our previous work, LEGCF, effectively reduced embedding storage costs but struggled to maintain recommendation performance under stricter storage limits. Additionally, LEGCF did not address the extensive runtime computation costs associated with graph propagation, which involves heavy multiplication and accumulation operations (MACs). These challenges consequently hinder effective training and inference on resource-constrained edge devices. To address these limitations, we propose Lightweight Embeddings with Rewired Graph for Graph Collaborative Filtering (LERG), an improved extension of LEGCF. LERG retains LEGCFs compositional codebook structure but introduces quantization techniques to reduce the storage cost, enabling the inclusion of more meta-embeddings within the same storage. To optimize graph propagation, we pretrain the quantized compositional embedding table using the full interaction graph on resource-rich servers, after which a fine-tuning stage is engaged to identify and prune low-contribution entities via a gradient-free binary integer programming approach, constructing a rewired graph that excludes these entities (i.e., user/item nodes) from propagating signals. The quantized compositional embedding table with selective embedding participation and sparse rewired graph are transferred to edge devices which significantly reduce computation memory and inference time. Experiments on three public benchmark datasets, including an industry-scale dataset, demonstrate that LERG achieves superior recommendation performance while dramatically reducing storage and computation costs for graph-based recommendation services.

Figures

Figures reproduced from arXiv: 2505.18999 by the authors.

Figure 1
Figure 1. The overall workflow of LERG. (a) corresponds to the quantized compositional embedding table [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. We organize our findings into the following subsections: [PITH_FULL_IMAGE:figures/full_fig_p018_2.png] view at source ↗
Figure 2
Figure 2. The plots on the left hand side show the recommendation performance of LERG w.r.t. different [PITH_FULL_IMAGE:figures/full_fig_p019_2.png] view at source ↗
Figures from the paper (2 more)
Figure 3
Figure 3. Figure 3: The performance of LERG w.r.t. various hyperparameter settings. [PITH_FULL_IMAGE:figures/full_fig_p022_3.png]
Figure 4
Figure 4. Figure 4: The performance of LERG when using different quantization precisions for [PITH_FULL_IMAGE:figures/full_fig_p024_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

114 extracted references · 75 canonical work pages

  1. [1]

    Yoshua Bengio, Nicholas Léonard, and Aaron Courville. 2013. Estimating or propagating gradients through stochastic neurons for conditional computation. (2013)

  2. [2]

    Pak K Chan, Martine DF Schlag, and Jason Y Zien. 1993. Spectral k-way ratio-cut partitioning and clustering. In Proceedings of the 30th international Design Automation Conference. 749–754

  3. [3]

    Jie Chen, Tengfei Ma, and Cao Xiao. 2018. Fastgcn: fast learning with graph convolutional networks via importance sampling. InICLR

  4. [4]

    Jianfei Chen, Jun Zhu, and Le Song. 2018. Stochastic Training of Graph Convolutional Networks with Variance Reduction. InICML

  5. [5]

    Lei Chen, Le Wu, Richang Hong, Kun Zhang, and Meng Wang. 2020. Revisiting graph based collaborative filtering: A linear residual graph convolutional network approach. InAAAI

  6. [6]

    Mengzhao Chen, Wenqi Shao, Peng Xu, Jiahao Wang, Peng Gao, Kaipeng Zhang, Yu Qiao, and Ping Luo. 2024. EfficientQAT: Efficient Quantization-Aware Training for Large Language Models. (2024)

  7. [7]

    Ming Chen, Zhewei Wei, Bolin Ding, Yaliang Li, Ye Yuan, Xiaoyong Du, and Ji-Rong Wen. 2020. Scalable graph neural networks via bidirectional propagation.NeurIPS(2020)

  8. [8]

    Tianlong Chen, Yongduo Sui, Xuxi Chen, Aston Zhang, and Zhangyang Wang. 2021. A unified lottery ticket hypothesis for graph neural networks. InICML

Show all 114 references
  1. [9]

    Tong Chen, Hongzhi Yin, Yujia Zheng, Zi Huang, Yang Wang, and Meng Wang. 2021. Learning elastic embeddings for customizing on-device recommenders. InKDD. 138–147

  2. [10]

    Wen Chen, Pipei Huang, Jiaming Xu, Xin Guo, Cheng Guo, Fei Sun, Chao Li, Andreas Pfadler, Huan Zhao, and Binqiang Zhao. 2019. POG: Personalized Outfit Generation for Fashion Recommendation at Alibaba iFashion. In KDD

  3. [11]

    Wei-Lin Chiang, Xuanqing Liu, Si Si, Yang Li, Samy Bengio, and Cho-Jui Hsieh. 2019. Cluster-gcn: An efficient algorithm for training deep and large graph convolutional networks. InKDD

  4. [12]

    Paul Covington, Jay Adams, and Emre Sargin. 2016. Deep neural networks for youtube recommendations. InRecsys

  5. [13]

    George B Dantzig. 1990. Origins of the simplex method. InA history of scientific computing. 141–151

  6. [14]

    Abhinandan S Das, Mayur Datar, Ashutosh Garg, and Shyam Rajaram. 2007. Google news personalization: scalable online collaborative filtering. InWWW

  7. [15]

    Aditya Desai, Yanzhou Pan, Kuangyuan Sun, Li Chou, and Anshumali Shrivastava. 2021. Semantically Constrained Memory Allocation (SCMA) for Embedding in Efficient Recommendation Systems. (2021)

  8. [16]

    Kaize Ding, Zhe Xu, Hanghang Tong, and Huan Liu. 2022. Data augmentation for deep graph learning: A survey. KDD24, 2 (2022)

  9. [17]

    DaYou Du, Yijia Zhang, Shijie Cao, Jiaqi Guo, Ting Cao, Xiaowen Chu, and Ningyi Xu. 2024. BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation. InACL. 102–116

  10. [18]

    Steven K Esser, Jeffrey L McKinstry, Deepika Bablani, Rathinakumar Appuswamy, and Dharmendra S Modha. 2020. LEARNED STEP SIZE QUANTIZATION. InICLR

  11. [19]

    Wenqi Fan, Yao Ma, Qing Li, Yuan He, Eric Zhao, Jiliang Tang, and Dawei Yin. 2019. Graph neural networks for social recommendation. InWWW

  12. [20]

    Chen Gao, Yu Zheng, Nian Li, Yinfeng Li, Yingrong Qin, Jinghua Piao, Yuhan Quan, Jianxin Chang, Depeng Jin, Xiangnan He, et al. 2023. A survey of graph neural networks for recommender systems: Challenges, methods, and directions.TORS1, 1 (2023), 1–51

  13. [21]

    Hui Guan, Andrey Malevich, Jiyan Yang, Jongsoo Park, and Hector Yuen. 2019. Post-training 4-bit quantization on embedding tables. (2019)

  14. [22]

    Vipul Gupta, Xin Chen, Ruoyun Huang, Fanlong Meng, Jianjun Chen, and Yujun Yan. 2024. GraphScale: A Framework to Enable Machine Learning over Billion-node Graphs. InCIKM

  15. [23]

    Saket Gurukar, Nikil Pancha, Andrew Zhai, Eric Kim, Samson Hu, Srinivasan Parthasarathy, Charles Rosenberg, and Jure Leskovec. 2022. MultiBiSage: A Web-Scale Recommendation System Using Multiple Bipartite Graphs at Pinterest. VLDB16, 4 (2022), 781–789

  16. [24]

    Will Hamilton, Zhitao Ying, and Jure Leskovec. 2017. Inductive representation learning on large graphs.NeurIPS30 (2017)

  17. [25]

    Xiangnan He, Kuan Deng, Xiang Wang, Yan Li, YongDong Zhang, and Meng Wang. 2020. LightGCN: Simplifying and Powering Graph Convolution Network for Recommendation. InSIGIR. 639–648. , Vol. 1, No. 1, Article . Publication date: September 2018. Lightweight Embeddings with Graph Re...

  18. [26]

    Wenbing Huang, Tong Zhang, Yu Rong, and Junzhou Huang. 2018. Adaptive sampling towards fast graph representa- tion learning.NeurIPS31 (2018)

  19. [27]

    Joglekar, Cong Li, Mei Chen, Taibai Xu, Xiaoming Wang, Jay K

    Manas R. Joglekar, Cong Li, Mei Chen, Taibai Xu, Xiaoming Wang, Jay K. Adams, Pranav Khaitan, Jiahui Liu, and Quoc V. Le. 2020. Neural Input Search for Large Scale Recommendation Models. InKDD. 2387–2397

  20. [28]

    2009.Reducibility among combinatorial problems

    Richard M Karp. 2009.Reducibility among combinatorial problems. Springer. 219–241 pages

  21. [29]

    George Karypis and Vipin Kumar. 1997. METIS: A software package for partitioning unstructured graphs, partitioning meshes, and computing fill-reducing orderings of sparse matrices. (1997)

  22. [30]

    Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. (2014)

  23. [31]

    Thomas N Kipf and Max Welling. 2017. SEMI-SUPERVISED CLASSIFICATION WITH GRAPH CONVOLUTIONAL NETWORKS. InICLR

  24. [32]

    Pan Li, Maofei Que, and Alexander Tuzhilin. 2023. Dual contrastive learning for efficient static feature representation in sequential recommendations.TKDE36, 2 (2023), 544–555

  25. [33]

    Shiwei Li, Huifeng Guo, Lu Hou, Wei Zhang, Xing Tang, Ruiming Tang, Rui Zhang, and Ruixuan Li. 2023. Adaptive low-precision training for embeddings in click-through rate prediction. InAAAI, Vol. 37. 4435–4443

  26. [34]

    Shiwei Li, Huifeng Guo, Xing Tang, Ruiming Tang, Lu Hou, Ruixuan Li, and Rui Zhang. 2024. Embedding Compression in Recommender Systems: A Survey.Comput. Surveys56, 5 (2024), 1–21

  27. [35]

    Yang Li, Tong Chen, Peng-Fei Zhang, and Hongzhi Yin. 2021. Lightweight Self-Attentive Sequential Recommendation. InCIKM. 967–977

  28. [36]

    Defu Lian, Haoyu Wang, Zheng Liu, Jianxun Lian, Enhong Chen, and Xing Xie. 2020. LightRec: A Memory and Search-Efficient Recommender System. InWWW

  29. [37]

    Defu Lian, Xing Xie, Enhong Chen, and Hui Xiong. 2020. Product quantized collaborative filtering.TKDE33, 9 (2020)

  30. [38]

    Xurong Liang, Tong Chen, Lizhen Cui, Yang Wang, Meng Wang, and Hongzhi Yin. 2024. Lightweight Embeddings for Graph Collaborative Filtering. InSIGIR

  31. [39]

    Xurong Liang, Tong Chen, Quoc Viet Hung Nguyen, Jianxin Li, and Hongzhi Yin. 2023. Learning compact composi- tional embeddings via regularized pruning for recommendation. InICDM

  32. [40]

    Zhiqi Lin, Cheng Li, Youshan Miao, Yunxin Liu, and Yinlong Xu. 2020. Pagraph: Scaling gnn training on large graphs via computation-aware caching. InACM Symposium on Cloud Computing

  33. [41]

    Greg Linden, Brent Smith, and Jeremy York. 2003. Amazon. com recommendations: Item-to-item collaborative filtering.IEEE Internet computing7, 1 (2003), 76–80

  34. [42]

    Chong Liu, Defu Lian, Min Nie, and Xia Hu. 2020. Online optimized product quantization. InICDM

  35. [43]

    Haochen Liu, Xiangyu Zhao, Chong Wang, Xiaobing Liu, and Jiliang Tang. 2020. Automated Embedding Size Search in Deep Recommender Systems. InSIGIR

  36. [44]

    Qi Liu, Jin Zhang, Defu Lian, Yong Ge, Jianhui Ma, and Enhong Chen. 2021. Online additive quantization. InKDD

  37. [45]

    Ruixuan Liu, Yang Cao, Yanlin Wang, Lingjuan Lyu, Yun Chen, and Hong Chen. 2023. Privaterec: Differentially private model training and online serving for federated news recommendation. InKDD

  38. [46]

    Siyi Liu, Chen Gao, Yihong Chen, Depeng Jin, and Yong Li. 2021. Learnable Embedding Sizes for Recommender Systems.ICLR(2021)

  39. [47]

    Weiming Liu, Chaochao Chen, Xinting Liao, Mengling Hu, Jianwei Yin, Yanchao Tan, and Longfei Zheng. 2023. Federated Probabilistic Preference Distribution Modelling with Compactness Co-Clustering for Privacy-Preserving Multi-Domain Recommendation.. InIJCAI

  40. [48]

    Xin Liu, Mingyu Yan, Lei Deng, Guoqi Li, Xiaochun Ye, Dongrui Fan, Shirui Pan, and Yuan Xie. 2022. Survey on Graph Neural Network Acceleration: An Algorithmic Perspective. InIJCAI

  41. [49]

    Zechun Liu, Barlas Oguz, Changsheng Zhao, Ernie Chang, Pierre Stock, Yashar Mehdad, Yangyang Shi, Raghuraman Krishnamoorthi, and Vikas Chandra. 2024. LLM-QAT: Data-Free Quantization Aware Training for Large Language Models. InFindings of the ACL. 467–484

  42. [50]

    Stuart Lloyd. 1982. Least squares quantization in PCM.IEEE transactions on information theory28, 2 (1982), 129–137

  43. [51]

    Jing Long, Tong Chen, Quoc Viet Hung Nguyen, and Hongzhi Yin. 2023. Decentralized collaborative learning framework for next POI recommendation.TOIS41, 3 (2023), 1–25

  44. [52]

    Jing Long, Guanhua Ye, Tong Chen, Yang Wang, Meng Wang, and Hongzhi Yin. 2024. Diffusion-Based Cloud-Edge- Device Collaborative Learning for Next POI Recommendations. InKDD

  45. [53]

    Fuyuan Lyu, Xing Tang, Hong Zhu, Huifeng Guo, Yingxue Zhang, Ruiming Tang, and Xue Liu. 2022. OptEmbed: Learning Optimal Embedding Table for Click-through Rate Prediction. InCIKM. 1399–1409

  46. [54]

    Shuming Ma, Hongyu Wang, Lingxiao Ma, Lei Wang, Wenhui Wang, Shaohan Huang, Li Dong, Ruiping Wang, Jilong Xue, and Furu Wei. 2024. The era of 1-bit llms: All large language models are in 1.58 bits. (2024)

  47. [55]

    Nattaya Mairittha, Tittaya Mairittha, and Sozo Inoue. 2020. Improving activity data collection with on-device personalization using fine-tuning. InUbiComp/ISWC. , Vol. 1, No. 1, Article . Publication date: September 2018. 28 Xurong Liang et al

  48. [56]

    Kelong Mao, Jieming Zhu, Xi Xiao, Biao Lu, Zhaowei Wang, and Xiuqiang He. 2021. UltraGCN: ultra simplification of graph convolutional networks for recommendation. InCIKM

  49. [57]

    Markus Nagel, Marios Fournarakis, Rana Ali Amjad, Yelysei Bondarenko, Mart Van Baalen, and Tijmen Blankevoort

  50. [58]

    Aditya Pal, Chantat Eksombatchai, Yitong Zhou, Bo Zhao, Charles Rosenberg, and Jure Leskovec. 2020. Pinnersage: Multi-modal user embedding framework for recommendations at pinterest. InKDD

  51. [59]

    Liang Qu, Yonghong Ye, Ningzhi Tang, Lixin Zhang, Yuhui Shi, and Hongzhi Yin. 2022. Single-shot embedding dimension search in recommender system. InSIGIR

  52. [60]

    Yunke Qu, Tong Chen, Quoc Viet Hung Nguyen, and Hongzhi Yin. 2024. Budgeted embedding table for recommender systems. InWSDM

  53. [61]

    Yunke Qu, Tong Chen, Xiangyu Zhao, Lizhen Cui, Kai Zheng, and Hongzhi Yin. 2023. Continuous Input Embedding Size Search For Recommender Systems. InSIGIR

  54. [62]

    Yunke Qu, Liang Qu, Tong Chen, Xiangyu Zhao, Jianxin Li, and Hongzhi Yin. 2024. Sparser Training for On-Device Recommendation Systems. (2024)

  55. [63]

    Yunke Qu, Liang Qu, Tong Chen, Xiangyu Zhao, Quoc Viet Hung Nguyen, and Hongzhi Yin. 2024. Scalable Dynamic Embedding Size Search for Streaming Recommendation. InCIKM

  56. [64]

    Zheng Qu, Dimin Niu, Shuangchen Li, Hongzhong Zheng, and Yuan Xie. 2023. TT-GNN: Efficient On-Chip Graph Neural Network Training via Embedding Reformation and Hardware Optimization. InMICRO

  57. [65]

    Steffen Rendle. 2010. Factorization Machines. InICDM. 995–1000

  58. [66]

    Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme. 2009. BPR: Bayesian personalized ranking from implicit feedback. InUAI. 452–461

  59. [67]

    Yu Rong, Wenbing Huang, Tingyang Xu, and Junzhou Huang. 2020. Dropedge: Towards deep graph convolutional networks on node classification. InICLR

  60. [68]

    1998.Theory of linear and integer programming

    A Schrijver. 1998.Theory of linear and integer programming. John Wiley & Sons

  61. [69]

    Hao-Jun Michael Shi, Dheevatsa Mudigere, Maxim Naumov, and Jiyan Yang. 2020. Compositional Embeddings Using Complementary Partitions for Memory-Efficient Recommendation Systems.KDD(2020), 165–175

  62. [70]

    Jianing Sun, Zhaoyue Cheng, Saba Zuberi, Felipe Pérez, and Maksims Volkovs. 2021. Hgcf: Hyperbolic graph convolution networks for collaborative filtering. InWWW

  63. [71]

    Xiao Sun, Naigang Wang, Chia-Yu Chen, Jiamin Ni, Ankur Agrawal, Xiaodong Cui, Swagath Venkataramani, Kaoutar El Maghraoui, Vijayalakshmi Viji Srinivasan, and Kailash Gopalakrishnan. 2020. Ultra-low precision 4-bit training of deep neural networks.NeurIPS33 (2020)

  64. [72]

    Jake Topping, Francesco Di Giovanni, Benjamin Paul Chamberlain, Xiaowen Dong, and Michael M Bronstein. 2022. Understanding over-squashing and bottlenecks on graphs via curvature. InICLR

  65. [73]

    Hung Vinh Tran, Tong Chen, Nguyen Quoc Viet Hung, Zi Huang, Lizhen Cui, and Hongzhi Yin. 2025. A Thorough Performance Benchmarking on Lightweight Embedding-based Recommender Systems.TOIS(2025)

  66. [74]

    Hung Vinh Tran, Tong Chen, Guanhua Ye, Quoc Viet Hung Nguyen, Kai Zheng, and Hongzhi Yin. 2025. On-device content-based recommendation with single-shot embedding pruning: A cooperative game perspective. InWWW. 772–785

  67. [75]

    Ulrike Von Luxburg. 2007. A tutorial on spectral clustering.Statistics and computing17 (2007), 395–416

  68. [76]

    Zhongwei Wan, Xin Wang, Che Liu, Samiul Alam, Yu Zheng, Jiachen Liu, Zhongnan Qu, Shen Yan, Yi Zhu, Quanlu Zhang, et al. 2023. Efficient large language models: A survey.Transactions on Machine Learning Research(2023)

  69. [77]

    Hongyu Wang, Shuming Ma, Li Dong, Shaohan Huang, Huaijie Wang, Lingxiao Ma, Fan Yang, Ruiping Wang, Yi Wu, and Furu Wei. 2023. BitNet: Scaling 1-bit Transformers for Large Language Models. (2023)

  70. [78]

    Jizhe Wang, Pipei Huang, Huan Zhao, Zhibo Zhang, Binqiang Zhao, and Dik Lun Lee. 2018. Billion-scale commodity embedding for e-commerce recommendation in alibaba. InKDD

  71. [79]

    Qinyong Wang, Hongzhi Yin, Tong Chen, Zi Huang, Hao Wang, Yanchang Zhao, and Nguyen Quoc Viet Hung. 2020. Next Point-of-Interest Recommendation on Resource-Constrained Mobile Devices. InWWW. 906–916

  72. [80]

    Qinyong Wang, Hongzhi Yin, Tong Chen, Junliang Yu, Alexander Zhou, and Xiangliang Zhang. 2022. Fast-adapting and privacy-preserving federated recommender system.VLDB31, 5 (2022), 877–896

  73. [81]

    Xiang Wang, Xiangnan He, Meng Wang, Fuli Feng, and Tat-Seng Chua. 2019. Neural Graph Collaborative Filtering. InSIGIR. 165–174

  74. [82]

    Yining Wang, Liwei Wang, Yuanzhi Li, Di He, Wei Chen, and Tie-Yan Liu. 2013. A theoretical analysis of NDCG ranking measures. InCOLT, Vol. 8. 6

  75. [83]

    Yingchao Wang, Chen Yang, Shulin Lan, Liehuang Zhu, and Yan Zhang. 2024. End-edge-cloud collaborative computing for deep learning: A comprehensive survey.IEEE Communications Surveys & Tutorials(2024)

  76. [84]

    Kilian Weinberger, Anirban Dasgupta, John Langford, Alex Smola, and Josh Attenberg. 2009. Feature hashing for large scale multitask learning. InICML. , Vol. 1, No. 1, Article . Publication date: September 2018. Lightweight Embeddings with Graph Rewiring for Collaborative Filtering 29

  77. [85]

    Felix Wu, Amauri Souza, Tianyi Zhang, Christopher Fifty, Tao Yu, and Kilian Weinberger. 2019. Simplifying graph convolutional networks. InICML

  78. [86]

    Jiancan Wu, Xiang Wang, Fuli Feng, Xiangnan He, Liang Chen, Jianxun Lian, and Xing Xie. 2021. Self-supervised graph learning for recommendation. InSIGIR

  79. [87]

    Shiwen Wu, Fei Sun, Wentao Zhang, Xu Xie, and Bin Cui. 2022. Graph neural networks in recommender systems: a survey.Comput. Surveys55, 5 (2022), 1–37

  80. [88]

    Xiaoxia Wu, Cheng Li, Reza Yazdani Aminabadi, Zhewei Yao, and Yuxiong He. 2023. Understanding int4 quantization for language models: latency speedup, composability, and failure cases. InICML. 37524–37539

  81. [89]

    Haocheng Xi, Changhao Li, Jianfei Chen, and Jun Zhu. 2023. Training transformers with 4-bit integers.NeurIPS36 (2023)

  82. [90]

    Xin Xia, Hongzhi Yin, Junliang Yu, Qinyong Wang, Guandong Xu, and Quoc Viet Hung Nguyen. 2022. On-device next-item recommendation with self-supervised knowledge distillation. InSIGIR. 546–555

  83. [91]

    Xin Xia, Junliang Yu, Qinyong Wang, Chaoqun Yang, Nguyen Quoc Viet Hung, and Hongzhi Yin. 2023. Efficient on-device session-based recommendation.TOIS41, 4 (2023)

  84. [92]

    Xin Xia, Junliang Yu, Guandong Xu, and Hongzhi Yin. 2023. Towards communication-efficient model updating for on-device session-based recommendation. InCIKM

  85. [93]

    Zhiqiang Xu, Dong Li, Weijie Zhao, Xing Shen, Tianbo Huang, Xiaoyun Li, and Ping Li. 2021. Agile and accurate CTR prediction model training for massive-scale online advertising systems. InSIGMOD. 2404–2409

  86. [94]

    Yikai Yan, Chaoyue Niu, Renjie Gu, Fan Wu, Shaojie Tang, Lifeng Hua, Chengfei Lyu, and Guihai Chen. 2022. On-device learning for model personalization with large-scale cloud-coordinated domain adaption. InKDD

  87. [95]

    Liang Yang, Zesheng Kang, Xiaochun Cao, Di Jin 0001, Bo Yang, and Yuanfang Guo. 2019. Topology Optimization based Graph Convolutional Network.. InIJCAI

  88. [96]

    Jiangchao Yao, Feng Wang, Kunyang Jia, Bo Han, Jingren Zhou, and Hongxia Yang. 2021. Device-cloud collaborative learning for recommendation. InKDD

  89. [97]

    Yao Yao, Bin Liu, Haoxun He, Dakui Sheng, Ke Wang, Li Xiao, and Huanhuan Cao. 2023. i-Razor: A Differentiable Neural Input Razor for Feature Selection and Dimension Search in DNN-Based Recommender Systems.TKDE36, 9 (2023), 4736–4749

  90. [98]

    Zhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami, Jiali Yu, Eric Tan, Leyuan Wang, Qijing Huang, Yida Wang, Michael Mahoney, et al. 2021. Hawq-v3: Dyadic neural network quantization. InICML

  91. [99]

    Chunxing Yin, Bilge Acun, Carole-Jean Wu, and Xing Liu. 2021. TT-REC: Tensor train compression for deep learning recommendation model embeddings.Proceedings of Machine Learning and Systems3 (2021), 448–462

  92. [100]

    Chunxing Yin, Da Zheng, Israt Nisa, Christos Faloutsos, George Karypis, and Richard Vuduc. 2022. Nimble GNN Embedding with Tensor-Train Decomposition. InKDD. 2327–2335

  93. [101]

    Hongzhi Yin, Liang Qu, Tong Chen, Wei Yuan, Ruiqi Zheng, Jing Long, Xin Xia, Yuhui Shi, and Chengqi Zhang. 2024. On-device recommender systems: A comprehensive survey. (2024)

  94. [102]

    Junliang Yu, Xin Xia, Tong Chen, Lizhen Cui, Nguyen Quoc Viet Hung, and Hongzhi Yin. 2023. XSimGCL: Towards extremely simple graph contrastive learning for recommendation.TKDE36, 2 (2023)

  95. [103]

    Junliang Yu, Hongzhi Yin, Xin Xia, Tong Chen, Lizhen Cui, and Quoc Viet Hung Nguyen. 2022. Are graph augmenta- tions necessary? simple graph contrastive learning for recommendation. InSIGIR

  96. [104]

    Hanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan, and Viktor Prasanna. 2020. Graphsaint: Graph sampling based inductive learning method. InICLR

  97. [105]

    Caojin Zhang, Yicun Liu, Yuanpu Xie, Sofia Ira Ktena, Alykhan Tejani, Akshay Gupta, Pranay Kumar Myana, Deepak Dilipkumar, Suvadip Paul, Ikuhiro Ihara, et al. 2020. Model Size Reduction Using Frequency Based Double Hashing for Recommender Systems. InRecSys. 521–526

  98. [106]

    Chengyuan Zhang, Yang Wang, Lei Zhu, Jiayu Song, and Hongzhi Yin. 2021. Multi-graph heterogeneous interaction fusion for social recommendation.TOIS40, 2 (2021), 1–26

  99. [107]

    Honglei Zhang, Fangyuan Luo, Jun Wu, Xiangnan He, and Yidong Li. 2023. LightFR: Lightweight federated recom- mendation with privacy-preserving matrix factorization.TOIS41, 4 (2023)

  100. [108]

    Jin Zhang, Qi Liu, Defu Lian, Zheng Liu, Le Wu, and Enhong Chen. 2022. Anisotropic additive quantization for fast inner product search. InAAAI, Vol. 36

  101. [109]

    Xiangyu Zhao, Haochen Liu, Wenqi Fan, Hui Liu, Jiliang Tang, Chong Wang, Ming Chen, Xudong Zheng, Xiaobing Liu, and Xiwang Yang. 2021. Autoemb: Automated embedding dimensionality search in streaming recommendations. InICDM. 896–905

  102. [110]

    Cheng Zheng, Bo Zong, Wei Cheng, Dongjin Song, Jingchao Ni, Wenchao Yu, Haifeng Chen, and Wei Wang. 2020. Robust graph representation learning via neural sparsification. InICML

  103. [111]

    Ruiqi Zheng, Liang Qu, Tong Chen, Kai Zheng, Yuhui Shi, and Hongzhi Yin. 2024. Personalized elastic embedding learning for on-device recommendation.TKDE36, 7 (2024), 3363–3375. , Vol. 1, No. 1, Article . Publication date: September 2018. 30 Xurong Liang et al

  104. [112]

    Shangfei Zheng, Hongzhi Yin, Tong Chen, Quoc Viet Hung Nguyen, Wei Chen, and Lei Zhao. 2025. Do as I can, not as I get: Topology-aware multi-hop reasoning on multi-modal knowledge graphs.TKDE(2025)

  105. [113]

    Difan Zou, Ziniu Hu, Yewen Wang, Song Jiang, Yizhou Sun, and Quanquan Gu. 2019. Layer-dependent importance sampling for training deep and large graph convolutional networks.NeurIPS32 (2019). Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009 , Vol. 1, No. 1...

  106. [2021]

    A white paper on neural network quantization. (2021)

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.