Pith. sign in

Paper Citation Record · LEDGER

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM

As of 8 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2506.11108.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11108 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:52:37.467712Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 72d0b169-3be9-4ad3-b19a-a649f5e48bb2 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:37.384927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:37.384927Z digest=sha256:d724871e9715005815696eec5713644c4a474bc85962c4d0320ccf74a5251229

Observation 95fc398d-a297-4f1c-8d98-b458fc6fc4de · outbound

This paper cites Language models are few-shot learners.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Language models are few-shot learners

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:37.740477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:52:37.390175Z digest=sha256:8af1a2c553861a8a3489c8e6ccd25cfa2f9176fd0f57bec44e11393adf43b864

Observation 83ae9183-4e70-4620-b61b-65f555a32da2 · outbound

This paper cites an unresolved cited work.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:52:37.730864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:52:37.394048Z digest=sha256:f94fdf18816c5a1ca5a1976e60e40a43ade2763eb6318338f8c499d65e03905e

Observation f9e3cc82-2c77-42fe-9f70-a59af461c138 · outbound

This paper cites Proximal Policy Optimization Algorithms.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Proximal Policy Optimization Algorithms

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:37.397713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:37.397713Z digest=sha256:c4f572be8b5613f239b087fabeefe838501d2ddaf16ebc5cef3378acd7ec26fe

Observation c5296904-1e35-4486-90f4-70d32b9ce972 · outbound

This paper cites an unresolved cited work.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:52:37.721317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:52:37.401755Z digest=sha256:2389b6d2262d2cf40f0653d5b1f42a803089ab57f600e470d4c0ec333fe7043c

Observation e61d4da4-d49a-4d25-b718-a9518800205f · outbound

This paper cites Training language models to follow instructions with human feedback.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Training language models to follow instructions with human feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:37.405709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:37.405709Z digest=sha256:eaf4558eb08a2a498dde0de347562af64c441e822df11f45c705400d1eba7f8c

Observation ddc4387c-1470-43ea-9c0d-159fe149b04a · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:37.711382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:52:37.409857Z digest=sha256:93f4eb7e5fa44dfabcf0d4bd7a96240d05ce6be13749394ed06140e2e12ea6e5

Observation 0e4178c0-d51f-44b3-9909-0b9cf4367ed0 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Fine-Tuning Language Models from Human Preferences

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:37.413252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:37.413252Z digest=sha256:2f3fe8fc6fdffe85fed612924676b37a1ef584702c248beede2e4ed55e405990

Observation 69c3fb62-ea1b-4370-9c14-13645bb1e41c · outbound

This paper cites Forbidding Edges between Points in the Plane to Disconnect the Triangulation Flip Graph.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Forbidding Edges between Points in the Plane to Disconnect the Triangulation Flip Graph

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T05:52:37.537609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:52:37.417126Z digest=sha256:565338aa5f66c8d1f560de32d83ccf2fbe5a6dcd8c8ba7146fe2e0a722a9ba93

Observation 4ea29d57-e741-4927-bbac-f4f04f0a78e3 · outbound

This paper cites A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:37.420740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:37.420740Z digest=sha256:cd9b4136e6242b6ab424892f0809fb4682a1ac7e23c98254e091acf1499599e9

Observation 846290f6-a6b0-4630-b414-6dc7278644c8 · outbound

This paper cites Of Spiky SVDs and Music Recommendation.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Of Spiky SVDs and Music Recommendation

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T05:52:37.512806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:52:37.424314Z digest=sha256:714307bd49fd3d73c9d11ed966393b63cbc7fb07de8c2fcb30e85b9e6ae823ed

Observation c2cb8527-6812-487d-9d03-8a76b2428d1d · outbound

This paper cites MultiWOZ—a Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM MultiWOZ—a Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:37.701855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:52:37.427985Z digest=sha256:6c70ea31f7a12acd4019baf70014b7c1e24881a26302f24e94d246a71f4a19fc

Observation 4d91aee7-e62e-4922-b537-758cb1c5dc4e · outbound

This paper cites Revisiting Self-Training for Neural Sequence Production.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Revisiting Self-Training for Neural Sequence Production

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:37.692145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:52:37.431013Z digest=sha256:d5e5d277202fbcf3ae5cbd8244154e9c1d28a3697e27513d287ae8a285a3a89a

Observation 0a603438-15e4-4886-953d-5163a5e466af · outbound

This paper cites MSMatch: Semi-Supervised Multispectral Scene Classification with Few Labels.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM MSMatch: Semi-Supervised Multispectral Scene Classification with Few Labels

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:37.434235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:37.434235Z digest=sha256:96b12c3c8cb5ad37831614280d3c15b48e574a25c819477b48f48c835b684248

Observation 764905fd-a53d-4052-8763-0ddd164b93f6 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:37.682326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:52:37.437664Z digest=sha256:eb3f3df8cdc58d633aa345a4a71e0ce9bfa192c1be7320ef01123cf0390850dc

Observation 022d703c-b974-4e7b-93cf-0b867d6ea5ec · outbound

This paper cites Self-Coherence Score: Learning to Verify Reasoning Paths in Large Language Models.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Self-Coherence Score: Learning to Verify Reasoning Paths in Large Language Models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:37.672657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:52:37.440801Z digest=sha256:c05ad3311ae6751edabf3ac0b1b669db78e647cac096ce9ea5169cc7629ec151

Observation 8b29ce70-553b-44d5-95c6-95b31b4709a1 · outbound

This paper cites Towards Learning to Explain: An Attention-Based Self-supervised Approach for Improving Multi-turn Dialogue Mod- els.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Towards Learning to Explain: An Attention-Based Self-supervised Approach for Improving Multi-turn Dialogue Mod- els

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:37.662076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:52:37.443830Z digest=sha256:ce0e3627cdf89dce58585d1617f8b6fe703077b80aae7b62f3bca5af28b8e7ab

Observation eede9a13-b25e-4d6c-a7af-2b0f3b5ededd · outbound

This paper cites Deep Reinforcement Learning for Dialogue Generation.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Deep Reinforcement Learning for Dialogue Generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:37.651273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:52:37.447125Z digest=sha256:5ab29cae1cc0fbf303d99b789319c2a559b41955b20cdeb9e2f3f6af060d8483

Observation c4a4364f-c68c-4ec6-8331-d7f828067bec · outbound

This paper cites Serban, Michael Noseworthy, Laurent Charlin, Joelle Pineau.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Serban, Michael Noseworthy, Laurent Charlin, Joelle Pineau

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:37.640452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:52:37.450503Z digest=sha256:f613ccf1834ea19994df502596c52d9f14dd4e8f9ee33f3df287c74263dd05be

Observation d2a4a14f-8c88-460f-974c-d130dd9f710b · outbound

This paper cites Global Contextual Reinforcement Learning for Multi-Turn Dialogue.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Global Contextual Reinforcement Learning for Multi-Turn Dialogue

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:37.629095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:52:37.453810Z digest=sha256:d55bdd362f7168e4a6a1d877c8961d432228c8309c4dda5493a290d8a2445497

Observation 9e99d459-4187-4367-a9ff-e644468930e8 · outbound

This paper cites an unresolved cited work.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:52:37.618657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:52:37.457214Z digest=sha256:1fd47b9a691efa696cabc8a841314839746dfc3a9ad98b01e2242660bc2bc220

Observation f4ebcfa9-5b09-4cb7-a35d-de0961c06ae5 · outbound

This paper cites Self-Supervised Learning for Cross-Attention in Neural Text Generation.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Self-Supervised Learning for Cross-Attention in Neural Text Generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:37.608483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:52:37.460714Z digest=sha256:639c1a0d361cc9cc7b04334c2da586de52aa3b8e058205ea746bc6095f4e21b1

Observation eb3fbd76-b702-40c7-a15d-10a9e7fe4afc · outbound

This paper cites an unresolved cited work.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:52:37.597791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:52:37.464309Z digest=sha256:2e01a6eee8510fd3e648214c9d6707535e1b562501f8a6ab284c5078bfa83088

Observation a8349cd4-dc5c-4e84-8523-89a1cee13161 · outbound

This paper cites Ziegler, Ryan Lowe, Ilya Sutskever, and Paul F.

History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Ziegler, Ryan Lowe, Ilya Sutskever, and Paul F

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:52:37.587437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:52:37.467712Z digest=sha256:f080766b420e071198856e6297f4096f441c3496c803435ecbcdabca67dd39a0

Pith citing papers

No inbound Pith citation observations are available.