Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T17:47:43.553604Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2607.17524.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T17:47:43.553604Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3da58420-39b0-4edd-92a6-4a4ddfb8bae4 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Survey of hallucination in natural language generation.ACM computing surveys, 55(12):1–38,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41b6b21c-0365-4d8c-ac50-c5c6385ed9a1 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Understanding R1-Zero-Like Training: A Critical Perspective
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b061c23-3988-42e5-b43c-950eb3aa9a92 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Gradient Imbalance in Direct Preference Optimization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 842db8ad-6ec5-4f84-9d4f-5cc4f302b195 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbed606d-0c35-4571-98fd-1eacac1be6c2 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Fine-grained Hallucination Detection and Editing for Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0a87f47-dcf7-46dc-bbc3-a732970e3f33 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift doi: 10.1038/s41586-024-07335-x
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5817823-ff2c-4609-bd54-c90960077c26 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Proximal Policy Optimization Algorithms
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 688fed70-b9ff-4980-ad11-b4cc96599377 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaf141bd-b34e-4573-b83a-c96012071de3 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Understanding factual errors in summarization: Errors, summarizers, datasets, error detectors
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6114e41b-8180-477a-9910-8ae39f55d535 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Minicheck: Efficient fact-checking of llms on grounding documents
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86d193b3-c1ce-40bf-bf0e-65c57c5fc782 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Steering Language Models With Activation Engineering
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5a3c139-f196-4cae-aa08-cd8e0f2cc00a · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77cfbb00-f862-4196-adfc-dc88b6390736 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Neural Text Generation with Unlikelihood Training
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fe38bd6-a41b-4e3d-a2aa-0fcbddf74436 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Qwen3 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8419bea-dd29-4331-bf42-39254e5f7aab · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e631884d-355e-47d0-9e7c-6c823cd3d4cd · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Token-level Direct Preference Optimization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8069122d-7f13-4587-98ed-d1b5b9ffabb8 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Fine-Tuning Language Models from Human Preferences
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 909a909d-ff3e-4154-b870-0c392f81cb28 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Specifically, instead of decomposing the summary into individual sentences, we directly evaluate the full summary using the Bespoke-MiniCheck-7B model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46016e8d-6493-46e3-9c48-a24819a1be9c · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fd21c68-a7c1-4fba-89f4-aeb74e2991da · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift We report the total number of supervised tokens and the proportion of positive (factual) and negative (hallucinated) labels for each configuration
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3348cd36-1e17-4155-90d2-f7211246e2ce · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Spurious Rewards: Rethinking Training Signals in RLVR
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ad160cc-c056-47c0-8fff-665c7d7790fe · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0425af6-4c99-431e-b646-bc9bfe2fc42a · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Tldr: Token-level detective reward model for large vision language models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00bd0b4c-a756-45f3-9a03-77bb77936994 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Programming refusal with conditional ac- tivation steering
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93b194f0-782d-486a-90bb-8781124eb519 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift doi: 10.18653/v1/2024
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 841e4ae7-d903-4de7-98ee-5e41a1510015 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift The Llama 3 Herd of Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ee16ce7-4528-45de-8914-fca83a289185 · outbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Findings of the WMT 2024 shared task of the open language data initiative
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.