Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T20:21:20.974159Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 5 inbound Pith citation observations for arXiv:2502.10482.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T20:21:20.974159Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:52:37.420740Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T04:19:33.993409Z
13 of 13 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e3fafb5c-b76c-481f-a820-f06f2d865b78 · outbound
A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23a940bf-6709-4e65-a53b-f87b608dc08f · outbound
A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Language models are few-shot learners
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd570a5a-b346-4148-9799-631a8b5dd40f · outbound
A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5d495fb-a15a-4528-b2c4-72926c213062 · outbound
A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Chatgpt: Optimizing language models for dialogue
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 186287d1-4e2d-4e97-9dde-db2ae3999eb9 · outbound
A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Training language models to follow instructions with human feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19d814ff-2f57-4bdc-b3e7-45e508399abc · outbound
A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Language models are unsupervised multitask learn- ers
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e0af5e2c-0263-4138-926c-49e28a95f4b1 · outbound
A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Forbidding Edges between Points in the Plane to Disconnect the Triangulation Flip Graph
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a509ea76-56ad-464f-8786-0278026219ea · outbound
A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Proximal Policy Optimization Algorithms
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2cd1e53-ae47-4077-ba58-7ce3ddf7c156 · outbound
A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad398887-fec1-48f2-a2eb-38d9420a57f4 · outbound
A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Sutton and Andrew G
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffaf6ff3-7f47-4ed1-a8c6-423335e7a190 · outbound
A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Gomez, Łukasz Kaiser, and Illia Polosukhin
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d96af54b-3829-41ef-9387-7236d0200340 · outbound
A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e05ab74-bddc-45b4-bfe4-ddeaaa387a5a · outbound
A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Fine-Tuning Language Models from Human Preferences
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ea29d57-e741-4927-bbac-f4f04f0a78e3 · inbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89a34d12-678f-48f5-b3dd-6b2fa0dbdad5 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals
Reference 256
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 742f1105-7c47-4576-b0c5-e82d31d4c6d2 · inbound
Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9723710c-7b9f-495f-84c3-bd4a23e19024 · inbound
Trust Region On-Policy Distillation A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c1fe7b6d-0ea0-4f53-a2a8-6da86315c50a · inbound
When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.