Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:52:37.467712Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2506.11108.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:52:37.467712Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
24 of 24 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 72d0b169-3be9-4ad3-b19a-a649f5e48bb2 · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95fc398d-a297-4f1c-8d98-b458fc6fc4de · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Language models are few-shot learners
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83ae9183-4e70-4620-b61b-65f555a32da2 · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f9e3cc82-2c77-42fe-9f70-a59af461c138 · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Proximal Policy Optimization Algorithms
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5296904-1e35-4486-90f4-70d32b9ce972 · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e61d4da4-d49a-4d25-b718-a9518800205f · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Training language models to follow instructions with human feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddc4387c-1470-43ea-9c0d-159fe149b04a · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Gomez, Łukasz Kaiser, and Illia Polosukhin
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e4178c0-d51f-44b3-9909-0b9cf4367ed0 · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Fine-Tuning Language Models from Human Preferences
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69c3fb62-ea1b-4370-9c14-13645bb1e41c · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Forbidding Edges between Points in the Plane to Disconnect the Triangulation Flip Graph
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ea29d57-e741-4927-bbac-f4f04f0a78e3 · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 846290f6-a6b0-4630-b414-6dc7278644c8 · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Of Spiky SVDs and Music Recommendation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c2cb8527-6812-487d-9d03-8a76b2428d1d · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM MultiWOZ—a Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4d91aee7-e62e-4922-b537-758cb1c5dc4e · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Revisiting Self-Training for Neural Sequence Production
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a603438-15e4-4886-953d-5163a5e466af · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM MSMatch: Semi-Supervised Multispectral Scene Classification with Few Labels
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 764905fd-a53d-4052-8763-0ddd164b93f6 · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 022d703c-b974-4e7b-93cf-0b867d6ea5ec · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Self-Coherence Score: Learning to Verify Reasoning Paths in Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b29ce70-553b-44d5-95c6-95b31b4709a1 · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Towards Learning to Explain: An Attention-Based Self-supervised Approach for Improving Multi-turn Dialogue Mod- els
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eede9a13-b25e-4d6c-a7af-2b0f3b5ededd · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Deep Reinforcement Learning for Dialogue Generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4a4364f-c68c-4ec6-8331-d7f828067bec · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Serban, Michael Noseworthy, Laurent Charlin, Joelle Pineau
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d2a4a14f-8c88-460f-974c-d130dd9f710b · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Global Contextual Reinforcement Learning for Multi-Turn Dialogue
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9e99d459-4187-4367-a9ff-e644468930e8 · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f4ebcfa9-5b09-4cb7-a35d-de0961c06ae5 · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Self-Supervised Learning for Cross-Attention in Neural Text Generation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb3fbd76-b702-40c7-a15d-10a9e7fe4afc · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a8349cd4-dc5c-4e84-8523-89a1cee13161 · outbound
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM Ziegler, Ryan Lowe, Ilya Sutskever, and Paul F
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.