Pith. sign in

Paper Citation Record · LEDGER

xGR: Efficient Generative Recommendation Serving at Scale

As of 6 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2512.11529.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.11529 v3

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T16:54:18.900221Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T17:44:18.414311Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation c366178b-b35b-4ce5-a7cf-4afa4750c38e · outbound

This paper cites GPT-4 Technical Report.

xGR: Efficient Generative Recommendation Serving at Scale GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:14.776018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:14.776018Z digest=sha256:c64e77d1b573e7f066aa90c9b78cb531fa019bc5770a8608c5050dbd63a71827

Observation c569f7e9-ad95-4fc7-aa5f-a9aa5700409b · outbound

This paper cites Xgboost: A scalable tree boosting system.

xGR: Efficient Generative Recommendation Serving at Scale Xgboost: A scalable tree boosting system

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:14.839017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:14.839017Z digest=sha256:b3a73006e6da6ff2f1663a5c13ae412898df010f36d4a7954a2f8f1396f7912b

Observation 05cd57a5-f725-4ee2-bcdc-00049b6d75eb · outbound

This paper cites Wide & deep learning for recommender systems.

xGR: Efficient Generative Recommendation Serving at Scale Wide & deep learning for recommender systems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:14.909196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:14.909196Z digest=sha256:9dbd6969eed593f679b59d70579963f12be2bc388414bb06cf34409a84f40795

Observation 86c6b380-53a7-4216-931b-c4808d9f5e6b · outbound

This paper cites Deep neural networks for youtube recommendations.

xGR: Efficient Generative Recommendation Serving at Scale Deep neural networks for youtube recommendations

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:14.980894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:14.980894Z digest=sha256:edd5929fa6b20a5b6a454c3136ead8b782b9b083667a0e735ed28f31b63d8672

Observation 9d97c9a0-3ba3-4a90-a6a9-a62e942e2135 · outbound

This paper cites Flashattention: Fast and memory- efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344– 16359, 2022.

xGR: Efficient Generative Recommendation Serving at Scale Flashattention: Fast and memory- efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344– 16359, 2022

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:15.078468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:15.078468Z digest=sha256:78b1d145a49e63326f56d6cce4ed219a39c4d6ca790c7d00b978ab8d21937316

Observation 4cf09f93-ff46-4dd7-828b-e7947afea323 · outbound

This paper cites A review of modern recommender systems using generative models (gen-recsys).

xGR: Efficient Generative Recommendation Serving at Scale A review of modern recommender systems using generative models (gen-recsys)

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:15.148909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:15.148909Z digest=sha256:006d05a1a7ef72633b268bd8e0852b3fd0705d435507fd86f66e908b90acbdaa

Observation 1bdcf63e-eaa2-47b0-a7aa-da762e15418a · outbound

This paper cites OneRec: Unifying Retrieve and Rank with Generative Recommender and Iterative Preference Alignment.

xGR: Efficient Generative Recommendation Serving at Scale OneRec: Unifying Retrieve and Rank with Generative Recommender and Iterative Preference Alignment

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:15.211745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:15.211745Z digest=sha256:3675c60f5b3f238cfa9f7ed78bb30568ac02bfd37594af898ef506cfb38cd7af

Observation ce57243c-bb87-4c24-854b-d477197764c3 · outbound

This paper cites XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models.

xGR: Efficient Generative Recommendation Serving at Scale XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:15.310376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:15.310376Z digest=sha256:0d26f257a481f20e10703a1dc7f03edcdb1a77e9e4dfe516b2ee30ac5aefa70c

Observation 1cfd4f00-d144-459c-87ee-d9b3ea01dae2 · outbound

This paper cites Prefillonly: An infer- ence engine for prefill-only workloads in large language model applications.

xGR: Efficient Generative Recommendation Serving at Scale Prefillonly: An infer- ence engine for prefill-only workloads in large language model applications

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:15.409728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:15.409728Z digest=sha256:0784e47465aa076d062e415e5459a7593abc6d57db4e0a6e9bb8a87119cd85e7

Observation db1feced-e85e-43fb-bc59-9a62447306b3 · outbound

This paper cites Recommendation systems: Algorithms, challenges, metrics, and business opportunities.applied sciences, 10(21):7748, 2020.

xGR: Efficient Generative Recommendation Serving at Scale Recommendation systems: Algorithms, challenges, metrics, and business opportunities.applied sciences, 10(21):7748, 2020

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:15.508752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:15.508752Z digest=sha256:97669cadebfe9b8964802ebee1785d6e7c9352afbb85033d35e3041e82691938

Observation f5d64881-617f-495d-8d35-61492a106d85 · outbound

This paper cites Self-supervised learning on users’ spontaneous behaviors for multi-scenario ranking in e- commerce.

xGR: Efficient Generative Recommendation Serving at Scale Self-supervised learning on users’ spontaneous behaviors for multi-scenario ranking in e- commerce

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:15.629679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:15.629679Z digest=sha256:14584b2ac2b629734cb44553358cfc137a110b2c78f6a4ba13453844df4ab862

Observation 7b41b132-e123-425e-b046-627786003571 · outbound

This paper cites Gmlake: Efficient and transparent gpu memory defragmentation for large-scale dnn train- ing with virtual memory stitching.

xGR: Efficient Generative Recommendation Serving at Scale Gmlake: Efficient and transparent gpu memory defragmentation for large-scale dnn train- ing with virtual memory stitching

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:15.810865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:15.810865Z digest=sha256:d3ec09803d4cf83f14c3b52f62127947993aded81dbc20f6a94dec38828228ce

Observation aedf5704-43d9-4f40-8954-2927f9bfa854 · outbound

This paper cites DeepFM: A Factorization-Machine based Neural Network for CTR Prediction.

xGR: Efficient Generative Recommendation Serving at Scale DeepFM: A Factorization-Machine based Neural Network for CTR Prediction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:15.997576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:15.997576Z digest=sha256:6426c175421bfdf8389ba3ed4eb7d1e9a9348cbb5e142bb7bc20d07d1b34e393

Observation e5100d6f-9361-431e-b67e-eff45861ee3a · outbound

This paper cites Mtgr: Industrial-scale generative recommendation framework in meituan.

xGR: Efficient Generative Recommendation Serving at Scale Mtgr: Industrial-scale generative recommendation framework in meituan

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:16.137340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:16.137340Z digest=sha256:1764d238c36d18db9aa0913ecbb131980c4cfe28cf64a5d7f7b6bfb667e81f35

Observation 76b90a34-98da-44cf-b599-5d1daa47efad · outbound

This paper cites Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders.

xGR: Efficient Generative Recommendation Serving at Scale Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:16.289586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:16.289586Z digest=sha256:e6b7a5f745081a13be4a362d0b138de11e1dc808907ae9e04d68a3ca9734d51e

Observation ceaef617-cd0a-41b5-b83b-2e0b6b0592d6 · outbound

This paper cites Picture recommen- dation system built on instagram.

xGR: Efficient Generative Recommendation Serving at Scale Picture recommen- dation system built on instagram

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:16.380473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:16.380473Z digest=sha256:5834cf28fa4a190627a5e8e1a7b839188542bb5c77b2875b37016143239083e3

Observation 41be2315-1fa9-4fb4-b7b0-f850dbb928f2 · outbound

This paper cites Revisiting recom- mender systems: an investigative survey.Neural Com- puting and Applications, 37(4):2145–2173, 2025.

xGR: Efficient Generative Recommendation Serving at Scale Revisiting recom- mender systems: an investigative survey.Neural Com- puting and Applications, 37(4):2145–2173, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:16.460147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:16.460147Z digest=sha256:4f897cc53f42ecb07b619e8dab97fd307d6de61e588b4cbb0ba2e73c13da0609

Observation 3fcb98e7-db31-48b2-803a-3e857f96968c · outbound

This paper cites Scaling Laws for Neural Language Models.

xGR: Efficient Generative Recommendation Serving at Scale Scaling Laws for Neural Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:16.552005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:16.552005Z digest=sha256:4ab86c8e6622d3f0612461deaa2a1342cd1b9cac27f11b653e448a7f1ef9c694

Observation 4003f067-0f37-40fd-a33d-d5089744ef12 · outbound

This paper cites Hercules: Heterogeneity-aware inference serving for at-scale per- sonalized recommendation.

xGR: Efficient Generative Recommendation Serving at Scale Hercules: Heterogeneity-aware inference serving for at-scale per- sonalized recommendation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:16.630371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:16.630371Z digest=sha256:45182e380d51023ef3ea37d273d167b68ef4b8debf59f81346d0bcd17f158c57

Observation a015bb3c-6958-4996-a6be-9a17334e459c · outbound

This paper cites A survey of recommendation systems: recom- mendation models, techniques, and application fields.

xGR: Efficient Generative Recommendation Serving at Scale A survey of recommendation systems: recom- mendation models, techniques, and application fields

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:16.688644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:16.688644Z digest=sha256:2a8810bfb68e7a975f4ee4e3e6e42e06403cab4f0ad2d6520f70bc5af1b2a50d

Observation 3d76fa29-0808-415c-8c21-add7e0f4c757 · outbound

This paper cites Minionerec: An open-source framework for scaling generative recommendation.arXiv preprint arXiv:2510.24431, 2025.

xGR: Efficient Generative Recommendation Serving at Scale Minionerec: An open-source framework for scaling generative recommendation.arXiv preprint arXiv:2510.24431, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:16.776082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:16.776082Z digest=sha256:dbbd50057cc673bd0a7d1cd97dd4959e061e7e6d8409a4a5ae44368c8c8772cf

Observation e383b867-2ab9-4ba6-b7f1-9e5f6effc899 · outbound

This paper cites Efficient memory manage- ment for large language model serving with pagedatten- tion.

xGR: Efficient Generative Recommendation Serving at Scale Efficient memory manage- ment for large language model serving with pagedatten- tion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:16.841929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:16.841929Z digest=sha256:48c27dc71a3dd59c7633e3fd79ef4c758cb0db74dc875aeac7aaad0b9cbefde2

Observation cb9b6ace-37aa-4769-b9c7-ab3ae6bb14e3 · outbound

This paper cites Machine translation decoding beyond beam search.

xGR: Efficient Generative Recommendation Serving at Scale Machine translation decoding beyond beam search

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:16.943225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:16.943225Z digest=sha256:147d037f0598832fa9ca6f2424ce365407cc32c4b9497277f1a0bd472eada525

Observation aa61a01d-5197-461a-9a8f-178883f40c02 · outbound

This paper cites How can recommender systems benefit from large language models: A survey.ACM Transactions on Information Systems, 43(2):1–47, 2025.

xGR: Efficient Generative Recommendation Serving at Scale How can recommender systems benefit from large language models: A survey.ACM Transactions on Information Systems, 43(2):1–47, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:17.014368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:17.014368Z digest=sha256:e9d60e20ad86e4976381f3d9b5592d8f655ece6fca1d409fff6dc92d90a31cfc

Observation 25ffa2a8-2788-406d-a5bf-81ef8e2e62b9 · outbound

This paper cites Efficient Inference for Large Language Model-based Generative Recommendation.

xGR: Efficient Generative Recommendation Serving at Scale Efficient Inference for Large Language Model-based Generative Recommendation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:17.103527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:17.103527Z digest=sha256:056537a88c0b22a577203e9d1e39fbdd70b244e457e185bb87ae35f88760ffed

Observation 9b35df06-b711-4ec8-99ed-44f06639e7ef · outbound

This paper cites DeepSeek-V3 Technical Report.

xGR: Efficient Generative Recommendation Serving at Scale DeepSeek-V3 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:17.189844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:17.189844Z digest=sha256:f2deef41f4c4badfec9c24347fe9dfb58d3c443631d666319f9dd90514f83732

Observation 930e4476-837b-4159-a92e-3779fc6ac012 · outbound

This paper cites xllm technical report.

xGR: Efficient Generative Recommendation Serving at Scale xllm technical report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:17.290866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:17.290866Z digest=sha256:a4d87b04ee6fc0798974b21ad374b30efcb0e5964c6a532f61e9fb77cb5d0f4a

Observation e3accc95-0d03-4016-b8a0-d9bc660d249d · outbound

This paper cites Monolith: Real Time Recommendation System With Collisionless Embedding Table.

xGR: Efficient Generative Recommendation Serving at Scale Monolith: Real Time Recommendation System With Collisionless Embedding Table

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:17.370060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:17.370060Z digest=sha256:f3fa6ce82cd31e13636802a61a015f6826aaeb2f951e89a5e586fe5f5adf9548

Observation 825e3802-7d17-4154-a16b-e47a41cf5432 · outbound

This paper cites If beam search is the answer, what was the question? InPro- ceedings of the 2020 conference on empirical methods in natural language processing (emnlp), pages 2173–2185, 2020.

xGR: Efficient Generative Recommendation Serving at Scale If beam search is the answer, what was the question? InPro- ceedings of the 2020 conference on empirical methods in natural language processing (emnlp), pages 2173–2185, 2020

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:17.433614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:17.433614Z digest=sha256:b1009a60fb3fc5eb2baeceb3fe4c8ffa0c2b16640600a6b121b2ac85b5452d3d

Observation b25e41de-a802-4de9-aa35-38c01d869507 · outbound

This paper cites Deep Learning Recommendation Model for Personalization and Recommendation Systems.

xGR: Efficient Generative Recommendation Serving at Scale Deep Learning Recommendation Model for Personalization and Recommendation Systems

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:17.529002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:17.529002Z digest=sha256:304bc69a9a63b2703e7eaf04aa7f19770f799ee8f7d057b05c8bd40627499a3e

Observation ac4b3113-9000-4c6f-9cb6-c6c9f73bb853 · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting.

xGR: Efficient Generative Recommendation Serving at Scale Splitwise: Efficient generative llm inference using phase splitting

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:17.592116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:17.592116Z digest=sha256:266e53ba901399223f6bbada9caf6b44d858adce61c9507933db64af153a59db

Observation 66887b3d-1aa1-49e5-bf97-6e1673ffdc96 · outbound

This paper cites Recommender systems with generative retrieval.Ad- vances in Neural Information Processing Systems, 36:10299–10315, 2023.

xGR: Efficient Generative Recommendation Serving at Scale Recommender systems with generative retrieval.Ad- vances in Neural Information Processing Systems, 36:10299–10315, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:17.644324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:17.644324Z digest=sha256:13b75516861574222f9010f97de99ca8015869cad926b72bc9fff149a808bd34

Observation 4e204637-329e-483b-a9c9-18511a458a6a · outbound

This paper cites Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters.

xGR: Efficient Generative Recommendation Serving at Scale Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:17.726744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:17.726744Z digest=sha256:45e37b8080cd52995862b846893d4808192e45685ad01e107efd2feb0bcef827

Observation a2160b19-edd5-4b0a-be71-20032e0e7ac0 · outbound

This paper cites Ekko: A {Large- Scale} deep learning recommender system with {Low- Latency} model update.

xGR: Efficient Generative Recommendation Serving at Scale Ekko: A {Large- Scale} deep learning recommender system with {Low- Latency} model update

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:17.806626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:17.806626Z digest=sha256:3282f65a1a70c315f8088725ee7e3276b92715438280d51e7d39f77eb363de02

Observation 9d5c2b88-a4ab-45c3-bb09-23d960752ba2 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

xGR: Efficient Generative Recommendation Serving at Scale Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:17.862264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:17.862264Z digest=sha256:3199bdfefed81db66ac503210fab18e7cc86ed4af286f35785723bac5176b704

Observation fa3fe061-ac8f-4b5e-9ee2-d57c088d1973 · outbound

This paper cites Scaling transformers for discriminative recommendation via generative pre- training.

xGR: Efficient Generative Recommendation Serving at Scale Scaling transformers for discriminative recommendation via generative pre- training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:17.916072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:17.916072Z digest=sha256:5ec5b8609b922ead1bb6c6afbfb4ec6b3fe8dd760a09462cd65b0be3bb988a01

Observation 8b3141dd-07ee-4828-81af-547fc067a854 · outbound

This paper cites Atrec: Accelerating recommendation model train- ing on cpus.IEEE Transactions on Parallel and Dis- tributed Systems, 35(6):905–918, 2024.

xGR: Efficient Generative Recommendation Serving at Scale Atrec: Accelerating recommendation model train- ing on cpus.IEEE Transactions on Parallel and Dis- tributed Systems, 35(6):905–918, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:17.997774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:17.997774Z digest=sha256:f7b858d8853d791421001ff78ac93bf64d4263a10b8464c0f6a5b827905dba62

Observation ff5b75de-3aa1-4e58-86af-a707466f9035 · outbound

This paper cites A gpu-specialized inference parameter server for large-scale deep recom- mendation models.

xGR: Efficient Generative Recommendation Serving at Scale A gpu-specialized inference parameter server for large-scale deep recom- mendation models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:18.050114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:18.050114Z digest=sha256:88da02c90304bee45415151b35a432cdc1ff93f5cfcab7dbd42ac953e50e9578

Observation 7317d57b-b4ff-445b-8cd8-03d3f9f4fba4 · outbound

This paper cites Self- evaluation guided beam search for reasoning.Advances in Neural Information Processing Systems, 36:41618– 41650, 2023.

xGR: Efficient Generative Recommendation Serving at Scale Self- evaluation guided beam search for reasoning.Advances in Neural Information Processing Systems, 36:41618– 41650, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:18.103580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:18.103580Z digest=sha256:b656e23825856bd438ebe00d11edc854ef208993f678ef8e975aca0afe83dce2

Observation 782fb9b5-1823-43dd-a745-2f93667fd7cb · outbound

This paper cites Qwen3 Technical Report.

xGR: Efficient Generative Recommendation Serving at Scale Qwen3 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:18.150937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:18.150937Z digest=sha256:a5f4bcc725044b1c7a875e710e7c5e72b714ea7193b19481f25b9c0a9d93b709

Observation 04277e4a-992d-48c9-8777-23a0899fd1d1 · outbound

This paper cites {GPU-Disaggregated} serving for deep learning recommendation models at scale.

xGR: Efficient Generative Recommendation Serving at Scale {GPU-Disaggregated} serving for deep learning recommendation models at scale

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:18.202769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:18.202769Z digest=sha256:11a920fdfed2687d8267236a7ac3f2a3ef6d985d706f5810fd25476a34aed79f

Observation b4f4356c-7a93-47f5-8b92-2a9c662f6f32 · outbound

This paper cites Cacheblend: Fast large language model serving for rag with cached knowledge fusion.

xGR: Efficient Generative Recommendation Serving at Scale Cacheblend: Fast large language model serving for rag with cached knowledge fusion

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:18.254193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:18.254193Z digest=sha256:d68637fe37688252a0eb58b4953725b22a7506a27658a5944e12537eed016378

Observation 3e6c739b-de72-4f05-a3ff-717a57534e1e · outbound

This paper cites Grace: A scalable graph-based approach to accelerating recommendation model inference.

xGR: Efficient Generative Recommendation Serving at Scale Grace: A scalable graph-based approach to accelerating recommendation model inference

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:18.302107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:18.302107Z digest=sha256:b7fc503fb8edae54eb24ddffe753cb7278fc4a0e7b29120906f113e2ee5ac74d

Observation 88fe549c-fff7-412c-8f29-23602b0f4b81 · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based} generative models.

xGR: Efficient Generative Recommendation Serving at Scale Orca: A distributed serving system for {Transformer-Based} generative models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:18.360945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:18.360945Z digest=sha256:3cd3ea917dc62a1f953bc54d87d9cba4e2c687430dbf450164c4416c4c77aeb3

Observation fdf66428-9652-4e48-872c-d08f7750e71c · outbound

This paper cites Ic-cache: Efficient large language model serving via in-context caching.

xGR: Efficient Generative Recommendation Serving at Scale Ic-cache: Efficient large language model serving via in-context caching

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:18.427480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:18.427480Z digest=sha256:148508355b341aa5555ab1f3192a2bac3ced6a06357fe201a8a7a6b24dda4772

Observation d6616cb3-276b-4919-8751-a4ea021c2ff4 · outbound

This paper cites Evaluating recom- mender systems: survey and framework.ACM comput- ing surveys, 55(8):1–38, 2022.

xGR: Efficient Generative Recommendation Serving at Scale Evaluating recom- mender systems: survey and framework.ACM comput- ing surveys, 55(8):1–38, 2022

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:18.502942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:18.502942Z digest=sha256:746751fa81c6c196cec4cb02cb26f0ae38ba27aedccc9c4220edb130f82eb52b

Observation 5ede7056-9c76-4231-a11c-69d3855402e9 · outbound

This paper cites Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations.

xGR: Efficient Generative Recommendation Serving at Scale Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:18.584806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:18.584806Z digest=sha256:ebd7ee16101e71f5cf112184d566366dd7482a0693f039dda4c9da9a544d1f41

Observation a70c72ec-6bc7-456f-8837-c6501e096d51 · outbound

This paper cites Diffkv: Differentiated memory man- agement for large language models with parallel kv com- paction.

xGR: Efficient Generative Recommendation Serving at Scale Diffkv: Differentiated memory man- agement for large language models with parallel kv com- paction

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:18.667075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:18.667075Z digest=sha256:ab655feb23763a1b5893fbb63e227ef85a93604a29adda8a8c4472c260405eec

Observation 844255f3-a2c8-4300-8d72-ef17d98288d6 · outbound

This paper cites Sglang: Efficient execution of structured language model programs.Advances in neural information pro- cessing systems, 37:62557–62583, 2024.

xGR: Efficient Generative Recommendation Serving at Scale Sglang: Efficient execution of structured language model programs.Advances in neural information pro- cessing systems, 37:62557–62583, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:18.742390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:18.742390Z digest=sha256:e365a961e1943df8cf69e2ab5b4c6fde3810357111600b902cdc855c78fc8223

Observation b0e0300d-d120-4c8a-be6f-335d403cb4f1 · outbound

This paper cites {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving.

xGR: Efficient Generative Recommendation Serving at Scale {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:18.800449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:18.800449Z digest=sha256:f8dfdbf90d014c901c803b9d73d1d88a019b9cadd12488be9629626eeef7496f

Observation a48897a5-c25b-4940-8769-ac154fd2049e · outbound

This paper cites {NanoFlow}: Towards opti- mal large language model serving throughput.

xGR: Efficient Generative Recommendation Serving at Scale {NanoFlow}: Towards opti- mal large language model serving throughput

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:18.855454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:18.855454Z digest=sha256:d527c7862cba1d16addd336a7fa704a2e5ca09b520118dc5e74adabc716ab15d

Observation 1639c36f-62e9-450d-8f75-bd4dbc650227 · outbound

This paper cites Relayattention for efficient large language model serving with long system prompts.

xGR: Efficient Generative Recommendation Serving at Scale Relayattention for efficient large language model serving with long system prompts

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:18.900221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:18.900221Z digest=sha256:d7ce4ae797dedd86f0244352cdb95abbbed95a6e7c1a4305663f64cb0860aae1

Pith citing papers

Observation f43820cd-7a48-491e-b9ca-98325f5815bd · inbound

One Pool, Two Caches: Adaptive HBM Partitioning for Accelerating Generative Recommender Serving cites this paper.

One Pool, Two Caches: Adaptive HBM Partitioning for Accelerating Generative Recommender Serving xGR: Efficient Generative Recommendation Serving at Scale

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-30T02:15:36.626357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T17:44:18.414311Z digest=sha256:26b4c60b120faa6f41fc1d792753ac5b720a284e109409495e3aa993100f56df