Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:12:33.875901Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 3 inbound Pith citation observations for arXiv:2506.19235.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:12:33.875901Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T15:16:36.284072Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T17:25:51.272020Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 97ae766d-9af7-486d-8c84-23a3e1a74090 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34b0f6be-4581-4066-aff6-85c6a3977597 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Reinforcement learning based recommender systems: A survey
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 130a9e2c-3b63-4d38-9bf0-4fa19e436d94 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Language models are few-shot learners
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e760fceb-6a17-4112-99ae-838258a2de64 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Twin: Two-stage interest network for lifelong user behavior modeling in ctr prediction at kuaishou
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eed36b7d-a357-4938-98a8-6bf3c507493d · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Wide & deep learning for recommender systems
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24ee761f-9be1-44eb-8546-972628a2ca31 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Deep neural networks for youtube recommenda- tions
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7a2d09e1-dc40-4535-8cb9-a20d9ac91fbd · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Flashattention: Fast and memory-efficient exact attention with io-awareness
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc6abfcb-890a-4e8f-927e-4da5563e87b3 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Bert: Pre-training of deep bidirectional transformers for language understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c42f9214-9bfe-4dbc-8201-f17b69056f76 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28801878-da4f-4c7c-8d31-ec04558d3def · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 DeepFM: A Factorization-Machine based Neural Network for CTR Prediction
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2c7126d-f7b1-45eb-ae0a-6b06c28c8506 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea9f42bd-c89a-4cf1-b46f-25804ee9904f · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Neural collaborative filtering
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c2f5c2d-7d15-4d33-8493-b03ab0e29e96 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Session-based Recommendations with Recurrent Neural Networks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b11b211-b806-45c7-a53d-dc24d565891f · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Factual and Personalized Recommendations using Language Models and Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e9e8c6e-888d-46b0-ae30-8a298d9a3f66 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Genrec: Large language model for generative recommendation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f4f6a9e-d85c-4ac2-b6ca-0cb54f9f9c80 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Self-attentive sequential recommendation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 378042bc-9a5b-426e-ad55-61dc63f23aad · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Matrix factorization techniques for recom- mender systems
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3898939a-803c-4bef-b330-1e6d4fa7d26b · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 GPT4Rec: A Generative Framework for Personalized Recommendation and User Interests Interpretation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9daa28f9-c4e3-4621-b328-396bd34e7994 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Llara: Large language-recommendation assistant
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef15a54c-ffdd-406b-b832-fa79fd60956a · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Cascade ranking for operational e-commerce search
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ffb7c52d-7b33-4c9f-a3e8-666fd7c90c46 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Integrating large language models into recommendation via mutual augmentation and adaptive aggregation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f047dc14-655c-4a39-ac5f-a4da5245144b · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Human-level control through deep reinforcement learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d034671-8f2e-45a7-81f1-49609fe7a2db · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Training language models to follow instructions with human feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdef5190-5935-4a90-a143-b96c92849146 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Large language model based long-tail query rewriting in taobao search
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 057d0cc7-415f-4deb-83d2-6910b71dc400 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Rankflow: Joint optimization of multi-stage cascade ranking systems as flows
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e4ad6fc9-d1e7-44a4-9888-a0ac34df5051 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Direct preference optimization: Your language model is secretly a reward model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dbbb977-4b79-4568-8da1-7b7093d4630f · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Recommender systems with generative retrieval
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9f09756-f522-4026-be76-82782b589466 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Proximal Policy Optimization Algorithms
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3789c514-58a4-4d09-91aa-d06644610368 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d76c055-c758-4975-b431-e0b0da8ffbc5 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 HybridFlow: A Flexible and Efficient RLHF Framework
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f5341ee-59b9-4769-b59f-56cbbad49b78 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Learning to summarize with human feedback
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e3bc0a5-59c6-44f5-97d9-3cc99f7c03ef · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Bert4rec: Sequential recommendation with bidirectional encoder representations from transformer
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c78550a-e9ff-452b-bcbf-efc5fc1b878f · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 LLaMA: Open and Efficient Foundation Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 097aa803-1f42-4562-ab1e-f0891b819b15 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Kerl: A knowledge-guided reinforcement learning model for sequential recommendation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99a8b067-e3fc-4988-a874-df705a14f3f4 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Chain-of-thought prompting elicits reasoning in large language models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bd8d4bf-7194-45c8-b172-be0827d6c01a · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 PALR: Personalization Aware LLMs for Recommendation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 481ac885-beb2-4f72-a162-7174c3a6be73 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Feature-level deeper self-attention network for sequential recommendation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2cb6e1b4-a622-47e9-96fe-14893ee14302 · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 Deep interest network for click-through rate prediction
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37362ac1-d24e-4014-b74e-2b2e02f3b96a · outbound
RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1 S3-rec: Self-supervised learning for sequential recommendation with mutual information maximization
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13e11885-dd8a-4e4c-8915-2f8240635264 · inbound
RecRM-Bench: Benchmarking Multidimensional Reward Modeling for Agentic Recommender Systems RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ca45276-c3f9-447a-b9eb-192a5ea914a9 · inbound
Intuition-Guided Latent Reasoning for LLM-Based Recommendation RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 256a2bc7-e757-49fd-98e6-8408da8096ee · inbound
MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.