Pith. sign in

Paper Citation Record · LEDGER

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift

As of 17 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2607.17524.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.17524 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T17:47:43.553604Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3da58420-39b0-4edd-92a6-4a4ddfb8bae4 · outbound

This paper cites Survey of hallucination in natural language generation.ACM computing surveys, 55(12):1–38,.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Survey of hallucination in natural language generation.ACM computing surveys, 55(12):1–38,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:41.154463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:41.154463Z digest=sha256:0212cc948bd10677240721d24b50bd2fd93be68b282914e9c1bdab8cc197883b

Observation 41b6b21c-0365-4d8c-ac50-c5c6385ed9a1 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Understanding R1-Zero-Like Training: A Critical Perspective

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:41.428723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:41.428723Z digest=sha256:9b43a4be8a79c7bbf18a2957b38c3fffd9c266c96d4491185431eef1b39533ca

Observation 4b061c23-3988-42e5-b43c-950eb3aa9a92 · outbound

This paper cites Gradient Imbalance in Direct Preference Optimization.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Gradient Imbalance in Direct Preference Optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:41.496853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:41.496853Z digest=sha256:549240611ee5e929d86f71ab1e71dca37be36012bc473c1cbdf2ebb9f244b699

Observation 842db8ad-6ec5-4f84-9d4f-5cc4f302b195 · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:41.542364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:41.542364Z digest=sha256:a07bff55fafa0b17c52534ff0927863f86dd3b07b5176ee93c5be410bff40c8f

Observation fbed606d-0c35-4571-98fd-1eacac1be6c2 · outbound

This paper cites Fine-grained Hallucination Detection and Editing for Language Models.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Fine-grained Hallucination Detection and Editing for Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:41.610917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:41.610917Z digest=sha256:f64b7fc6b92360ce7286f712435a5d7c9714c81a98aced851c09b499ad51a86b

Observation a0a87f47-dcf7-46dc-bbc3-a732970e3f33 · outbound

This paper cites doi: 10.1038/s41586-024-07335-x.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift doi: 10.1038/s41586-024-07335-x

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:41.699281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:41.699281Z digest=sha256:37f4ac2e0ff43963c951af6734033d8633f66e24b475930eb495d071e0d29d2a

Observation a5817823-ff2c-4609-bd54-c90960077c26 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Proximal Policy Optimization Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:41.779404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:41.779404Z digest=sha256:1233e971f44787874d2b05d0b03a0e991f90038c1f1648ce42fd2e9e14e2174b

Observation 688fed70-b9ff-4980-ad11-b4cc96599377 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:41.914703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:41.914703Z digest=sha256:93070f9f593ab9e147842586184db2c0c391efdd067181b240c4458e79381046

Observation eaf141bd-b34e-4573-b83a-c96012071de3 · outbound

This paper cites Understanding factual errors in summarization: Errors, summarizers, datasets, error detectors.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Understanding factual errors in summarization: Errors, summarizers, datasets, error detectors

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:41.994642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:41.994642Z digest=sha256:875dc62a3e19630e27601a0fe243b5b9529a0c786eb22291848a6c403c61ecf6

Observation 6114e41b-8180-477a-9910-8ae39f55d535 · outbound

This paper cites Minicheck: Efficient fact-checking of llms on grounding documents.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Minicheck: Efficient fact-checking of llms on grounding documents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:42.112835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:42.112835Z digest=sha256:a54dc75c5cdd24b1d8a0c2c7477cfeb0db31ac685114fbc57071e239842c7457

Observation 86d193b3-c1ce-40bf-bf0e-65c57c5fc782 · outbound

This paper cites Steering Language Models With Activation Engineering.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Steering Language Models With Activation Engineering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:42.254879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:42.254879Z digest=sha256:e68811741ade3c7551fafad7238cd61051de224496f57e5038a7d63ef18b638f

Observation c5a3c139-f196-4cae-aa08-cd8e0f2cc00a · outbound

This paper cites LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:42.427244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:42.427244Z digest=sha256:c1e5d3837f7bd2f171380524977c4dd28d78c1bba37ec8dfe2eb4df15d7be1d4

Observation 77cfbb00-f862-4196-adfc-dc88b6390736 · outbound

This paper cites Neural Text Generation with Unlikelihood Training.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Neural Text Generation with Unlikelihood Training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:42.511417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:42.511417Z digest=sha256:94a1e8a467195d99338a9464a9eebe5b2bc55969f13601e97ac12d8e7598aa73

Observation 3fe38bd6-a41b-4e3d-a2aa-0fcbddf74436 · outbound

This paper cites Qwen3 Technical Report.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Qwen3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:42.664555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:42.664555Z digest=sha256:0ad4949ecf3b9747804219a63c1a0bf29ee2386ad80fca19f1b55d5a1b9757ff

Observation c8419bea-dd29-4331-bf42-39254e5f7aab · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:42.786853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:42.786853Z digest=sha256:16756bb54c7d9a06a75cd668654232df99fa41cc873e648c8155b3310ff08b83

Observation e631884d-355e-47d0-9e7c-6c823cd3d4cd · outbound

This paper cites Token-level Direct Preference Optimization.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Token-level Direct Preference Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:42.921201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:42.921201Z digest=sha256:580245009223e1d81578f101c96cd32fce8f2b3eaf1d6c1a95958c6998242e80

Observation 8069122d-7f13-4587-98ed-d1b5b9ffabb8 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Fine-Tuning Language Models from Human Preferences

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:43.057405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:43.057405Z digest=sha256:31bf154a93e982d9940b50680c88d2ddf00069505fc4219ba5a49e1c91d6d606

Observation 909a909d-ff3e-4154-b870-0c392f81cb28 · outbound

This paper cites Specifically, instead of decomposing the summary into individual sentences, we directly evaluate the full summary using the Bespoke-MiniCheck-7B model.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Specifically, instead of decomposing the summary into individual sentences, we directly evaluate the full summary using the Bespoke-MiniCheck-7B model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:43.165843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:43.165843Z digest=sha256:5cb79cb1871814ff9660f2ab243731308fe74455b9b96d5b7030ff99354a7821

Observation 46016e8d-6493-46e3-9c48-a24819a1be9c · outbound

This paper cites an unresolved cited work.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:43.343218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:43.343218Z digest=sha256:b1d60934913b1f98ffc37dbea549c07cd04163a310f037ef5386159bfb81d6ef

Observation 6fd21c68-a7c1-4fba-89f4-aeb74e2991da · outbound

This paper cites We report the total number of supervised tokens and the proportion of positive (factual) and negative (hallucinated) labels for each configuration.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift We report the total number of supervised tokens and the proportion of positive (factual) and negative (hallucinated) labels for each configuration

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:43.553604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:43.553604Z digest=sha256:99b33a43eff4d53341240d9e134ab832aa16c9ccca3183e241797cfe9bca0996

Observation 3348cd36-1e17-4155-90d2-f7211246e2ce · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Spurious Rewards: Rethinking Training Signals in RLVR

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:41.835722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:41.835722Z digest=sha256:95e65cc01a5abf86caccd306aa1749a7fba2e8b82f12b6111615c87b24074896

Observation 2ad160cc-c056-47c0-8fff-665c7d7790fe · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:42.557783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:42.557783Z digest=sha256:4b4f804d870eb4f4502bb77a1d68a130fab0009d30211b8a7ba95fc4e6cac055

Observation e0425af6-4c99-431e-b646-bc9bfe2fc42a · outbound

This paper cites Tldr: Token-level detective reward model for large vision language models.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Tldr: Token-level detective reward model for large vision language models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:40.966435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:40.966435Z digest=sha256:06c0abf9860fe9cb615044bf2e826ec7089277fb1b94351806ea189a9eb95ebe

Observation 00bd0b4c-a756-45f3-9a03-77bb77936994 · outbound

This paper cites Programming refusal with conditional ac- tivation steering.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Programming refusal with conditional ac- tivation steering

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:41.295760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:41.295760Z digest=sha256:4a64ce68c0d038bde3bf7f9e91fda18b233c52f6953f6a62edf1b13eac13a223

Observation 93b194f0-782d-486a-90bb-8781124eb519 · outbound

This paper cites doi: 10.18653/v1/2024.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift doi: 10.18653/v1/2024

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:40.898240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:40.898240Z digest=sha256:a7bf536edec25cdfd4261c0bb58fa109c26bc71a385002336dda8ab9df0e442a

Observation 841e4ae7-d903-4de7-98ee-5e41a1510015 · outbound

This paper cites The Llama 3 Herd of Models.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift The Llama 3 Herd of Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:41.056047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:41.056047Z digest=sha256:6aff79820e11394dc91747f9b73ffaa5796db48a9ffb5ffe0a6ac636126c7e50

Observation 7ee16ce7-4528-45de-8914-fca83a289185 · outbound

This paper cites Findings of the WMT 2024 shared task of the open language data initiative.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Findings of the WMT 2024 shared task of the open language data initiative

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:40.833651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:40.833651Z digest=sha256:af14867ab6a5075f37ee132408d53b09795bddcf38dab4bb65a747150d3a04bb

Pith citing papers

No inbound Pith citation observations are available.