Pith. sign in

Paper Citation Record · LEDGER

A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals

As of 9 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 5 inbound Pith citation observations for arXiv:2502.10482.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.10482 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:21:20.974159Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:52:37.420740Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T04:19:33.993409Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e3fafb5c-b76c-481f-a820-f06f2d865b78 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:20.913663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:20.913663Z digest=sha256:53549865c9791b446d77817af04043bc10c965a848b59446007450db76d32002

Observation 23a940bf-6709-4e65-a53b-f87b608dc08f · outbound

This paper cites Language models are few-shot learners.

A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Language models are few-shot learners

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:20.920048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:20.920048Z digest=sha256:06ba3edb31936e2a5eee46e878ec6f7816745fbdf7105f2597098da24cebe489

Observation cd570a5a-b346-4148-9799-631a8b5dd40f · outbound

This paper cites an unresolved cited work.

A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:20.924816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:20.924816Z digest=sha256:7a97b5afd7d98036d5616754f22763ffad11ba1aa96ca5a14e4f703847f5ba48

Observation b5d495fb-a15a-4528-b2c4-72926c213062 · outbound

This paper cites Chatgpt: Optimizing language models for dialogue.

A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Chatgpt: Optimizing language models for dialogue

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:21:21.181148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T20:21:20.929625Z digest=sha256:11009649cfee731651c3195756b26a3924c74757cb0ec04cee528ff951eba85d

Observation 186287d1-4e2d-4e97-9dde-db2ae3999eb9 · outbound

This paper cites Training language models to follow instructions with human feedback.

A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Training language models to follow instructions with human feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:20.934897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:20.934897Z digest=sha256:0c4429950a930b3d006688b8b14963ab44de030ccc5b27f712461edd807bcc30

Observation 19d814ff-2f57-4bdc-b3e7-45e508399abc · outbound

This paper cites Language models are unsupervised multitask learn- ers.

A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Language models are unsupervised multitask learn- ers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:21:21.162171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T20:21:20.939914Z digest=sha256:0b4db620b66fae737d52c058f9b99ca17a256d051444f5c927c1431349678520

Observation e0af5e2c-0263-4138-926c-49e28a95f4b1 · outbound

This paper cites Forbidding Edges between Points in the Plane to Disconnect the Triangulation Flip Graph.

A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Forbidding Edges between Points in the Plane to Disconnect the Triangulation Flip Graph

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:20.945097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:20.945097Z digest=sha256:14c1f3a136511cf4e5686f37fa46bf690a3019a8b68e82d6a35b070cbf5c62a1

Observation a509ea76-56ad-464f-8786-0278026219ea · outbound

This paper cites Proximal Policy Optimization Algorithms.

A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Proximal Policy Optimization Algorithms

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:20.949944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:20.949944Z digest=sha256:c0824242ac7e2b48d068ebff1611cdffd49909db1c755df3fd2ec7a00c8132bd

Observation c2cd1e53-ae47-4077-ba58-7ce3ddf7c156 · outbound

This paper cites an unresolved cited work.

A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:20.954926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:20.954926Z digest=sha256:326bbef30df21470a7741dd35092012511a663c009cd796db279585224c7c4c8

Observation ad398887-fec1-48f2-a2eb-38d9420a57f4 · outbound

This paper cites Sutton and Andrew G.

A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Sutton and Andrew G

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:20.959735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:20.959735Z digest=sha256:735d5105930f72752ef5d953062271f395a3040f40b01ba0dd2836ecfaeb6138

Observation ffaf6ff3-7f47-4ed1-a8c6-423335e7a190 · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:20.964290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:20.964290Z digest=sha256:b3bcc03c9008d94adc27b40fa083dc55fecb29c9401fabb26b8feffeb15a0a55

Observation d96af54b-3829-41ef-9387-7236d0200340 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:20.969004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:20.969004Z digest=sha256:ae600e914360cc5a0154902598d011b9518e6ab71c900c04c0196d5d0e0dafa1

Observation 7e05ab74-bddc-45b4-bfe4-ddeaaa387a5a · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Fine-Tuning Language Models from Human Preferences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:20.974159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:20.974159Z digest=sha256:f929a10bc1aabe390ed8eaa9eceb991dd9014f55181c19ef5d31229b37896b68

Pith citing papers

Observation 4ea29d57-e741-4927-bbac-f4f04f0a78e3 · inbound

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM cites this paper.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:37.420740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:37.420740Z digest=sha256:cd9b4136e6242b6ab424892f0809fb4682a1ac7e23c98254e091acf1499599e9

Observation 89a34d12-678f-48f5-b3dd-6b2fa0dbdad5 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals

Reference 256

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:24.643178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:5264bcd9b4e51e6bd189b0ecb2992ccfd8679deb7c2e110a4535bad1413dc2b1

Observation 742f1105-7c47-4576-b0c5-e82d31d4c6d2 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.825632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.825632Z digest=sha256:62a165a3adeb37a38386ad41e0211324b9db6c12e7fede12329763d98456964a

Observation 9723710c-7b9f-495f-84c3-bd4a23e19024 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.494301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:701c878e89325ad97753ab492ff5c0a4d5b880864d89e2dd3287ba410f3fade8

Observation c1fe7b6d-0ea0-4f53-a2a8-6da86315c50a · inbound

When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study cites this paper.

When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:19:33.996029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T17:07:21.486960Z digest=sha256:f0ac1dfe36257a3d0ded1fcf68120f9886bca3efcef937fcb62de767318c4b89