Pith. sign in

Paper Citation Record · LEDGER

Lessons from Training Grounded LLMs with Verifiable Rewards

As of 21 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 3 inbound Pith citation observations for arXiv:2506.15522.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15522 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:59:49.564233Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:48:29.575024Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:09:47.204727Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 35225b90-78d7-4752-be6b-ec1e6cd2e9f6 · outbound

This paper cites I apologize, but I couldn't find an answer to your question in the search results.

Lessons from Training Grounded LLMs with Verifiable Rewards I apologize, but I couldn't find an answer to your question in the search results

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:59:50.547041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:59:49.564233Z digest=sha256:d4b079dddb8e8bd34c0a721fd3762d92713f2ef384036eb77e6ee18bc4027c7d

Observation 0230525b-aa15-4555-b9e9-234d6bde6b96 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Lessons from Training Grounded LLMs with Verifiable Rewards DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.123573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.123573Z digest=sha256:5c3d85829b4483c83b520fde659d1f11baddc69c28830e3e264d8c0892bc6937

Observation 45617367-ef31-4fb2-9d91-6a7148d34bca · outbound

This paper cites Training Language Models to Generate Text with Citations via Fine-grained Rewards.

Lessons from Training Grounded LLMs with Verifiable Rewards Training Language Models to Generate Text with Citations via Fine-grained Rewards

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.358493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.358493Z digest=sha256:717163d1f98d202134e959b842ac430fcfd5ccb1a7c6841e75db2f7500675d1f

Observation 3daa139d-305f-433a-ab13-407f4c4408fc · outbound

This paper cites RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement.

Lessons from Training Grounded LLMs with Verifiable Rewards RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.438690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.438690Z digest=sha256:49f51c38124d75956e310937e00effb57b9a43d7a18d2cb15eb8879ef01fd35d

Observation 300387b0-b5d7-485e-a35a-6b66106dc482 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Lessons from Training Grounded LLMs with Verifiable Rewards Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.887449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.887449Z digest=sha256:9e88f2f77a5d5b8b7beccb3fd894fc1ebec1d285bb6fd45a3035a61993ff3a29

Observation a7343cb3-fb7d-42ca-b9e2-6bbd434ee48e · outbound

This paper cites Attribute First, then Generate: Locally-attributable Grounded Text Generation.

Lessons from Training Grounded LLMs with Verifiable Rewards Attribute First, then Generate: Locally-attributable Grounded Text Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.945371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.945371Z digest=sha256:c6040e70714fe807a3b70b2070baf4023e74ece3ce5b2f8901514de8440dabd6

Observation 38d883f4-49ef-47d8-9173-84b1fcecba0e · outbound

This paper cites Qwen3 Technical Report.

Lessons from Training Grounded LLMs with Verifiable Rewards Qwen3 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:49.122177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:49.122177Z digest=sha256:96baf6c09f2ca99d20abd1cd02f7a4447505089b12faba251e53a08f41766af3

Observation 9a973654-1c0c-499d-bdfe-82e746fde4a2 · outbound

This paper cites Ground Every Sentence: Improving Retrieval-Augmented LLMs with Interleaved Reference-Claim Generation.

Lessons from Training Grounded LLMs with Verifiable Rewards Ground Every Sentence: Improving Retrieval-Augmented LLMs with Interleaved Reference-Claim Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:49.203531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:49.203531Z digest=sha256:85cd4ba46fa8bfd08749e82ecc09cf165ca463426087f4454afc39b22d18de8e

Observation bc6b8e4e-2b09-48c3-8f94-c45a11e00523 · outbound

This paper cites RECOMP: Improving Retrieval-Augmented LMs with Compression and Selective Augmentation.

Lessons from Training Grounded LLMs with Verifiable Rewards RECOMP: Improving Retrieval-Augmented LMs with Compression and Selective Augmentation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:49.286759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:49.286759Z digest=sha256:d181dc6f397aaeb5a71ff683bfd2ed017b8775c535adc24a32ef8973e3d71afc

Observation 18acf74a-00a1-4c3b-9ccc-2daf31e15d24 · outbound

This paper cites Effective Large Language Model Adaptation for Improved Grounding and Citation Generation.

Lessons from Training Grounded LLMs with Verifiable Rewards Effective Large Language Model Adaptation for Improved Grounding and Citation Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:49.340475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:49.340475Z digest=sha256:75410ed84a94535a780fb78bfafb3ac4a5e5eeaacb6c625f64088e2d7dbe219c

Observation 125741e5-9e1a-47bc-8313-69b1c8f423f7 · outbound

This paper cites Making Retrieval-Augmented Language Models Robust to Irrelevant Context.

Lessons from Training Grounded LLMs with Verifiable Rewards Making Retrieval-Augmented Language Models Robust to Irrelevant Context

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:49.441945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:49.441945Z digest=sha256:8d4095c80b929fe2da03a113b1750c3f113468f13b2bfe49ea00002181e828cb

Observation d956c89f-400c-4804-9a74-e5be51119b0c · outbound

This paper cites In Webber, B.; Cohn, T.; He, Y .; and Liu, Y ., eds.,Proceedings of the 2020 Confer- ence on Empirical Methods in Natural Language Processing (EMNLP), 6769–6781.

Lessons from Training Grounded LLMs with Verifiable Rewards In Webber, B.; Cohn, T.; He, Y .; and Liu, Y ., eds.,Proceedings of the 2020 Confer- ence on Empirical Methods in Natural Language Processing (EMNLP), 6769–6781

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:59:50.767047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:59:48.580786Z digest=sha256:42225443f982269cc1eb99237838368b13bb716cc8a881c0b815ba89ae027e1d

Observation ee134af0-4c9d-4123-91d5-14d8f7f6ab1d · outbound

This paper cites Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

Lessons from Training Grounded LLMs with Verifiable Rewards Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.657029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.657029Z digest=sha256:68b5ab5191f54d7ed39b399bd2ab71f0ad3ad030740f63fe72009f62629b4023

Observation 9c66798b-d9aa-4d47-8cd1-81052fcb18f4 · outbound

This paper cites Training language models to follow instructions with human feedback.

Lessons from Training Grounded LLMs with Verifiable Rewards Training language models to follow instructions with human feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.756238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.756238Z digest=sha256:1a63d0212f4f5b42ed777b5b8da3076a964ed15219b31cb1381df80b7d433e0a

Observation 96fcc4ea-7249-48d0-9135-ca1ca25c81ce · outbound

This paper cites Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection.

Lessons from Training Grounded LLMs with Verifiable Rewards Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:47.915712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:47.915712Z digest=sha256:36ebb7186a9975cd2615dd3d49c3997f585bff48dd8243dcd2e62d30cc3b9af4

Observation 9fe83f6a-6464-4523-bb30-8bd7adb02a88 · outbound

This paper cites The Llama 3 Herd of Models.

Lessons from Training Grounded LLMs with Verifiable Rewards The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.240829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.240829Z digest=sha256:9e7963fa4e96aa1d4bef6067c2559a78fbc022b0c9b778d87e7cb3c84c2abdd5

Observation fb49ac51-e732-439e-9968-a068c4997491 · outbound

This paper cites ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning.

Lessons from Training Grounded LLMs with Verifiable Rewards ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:47.976127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:47.976127Z digest=sha256:8b90182e6d269ee911d6e26730f1da681355efc7150a1b2da2f2792dde910f05

Pith citing papers

Observation 49264500-90d5-4536-82ab-ac3512717904 · inbound

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards cites this paper.

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards Lessons from Training Grounded LLMs with Verifiable Rewards

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:26:28.297833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T14:24:48.666197Z digest=sha256:06a49f0ae9105c1352d17d50545a9430a461e7f280cc5012627d15a3f2cff6b5

Observation 3f8a0b7a-1887-4d68-8f7d-2f074cdb4939 · inbound

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards cites this paper.

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards Lessons from Training Grounded LLMs with Verifiable Rewards

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T15:48:29.575024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:48:29.575024Z digest=sha256:19ce53d12d9f51e5cb7df544da23a624b89b2c831c43250a70d6e5c8cee60d34

Observation dc08d080-03c2-41cf-88a2-737c511970a8 · inbound

Do LLM Attribution Metrics Transfer? Auditing Retrieval-Augmented Generation Evaluation Across Datasets and Constructs cites this paper.

Do LLM Attribution Metrics Transfer? Auditing Retrieval-Augmented Generation Evaluation Across Datasets and Constructs Lessons from Training Grounded LLMs with Verifiable Rewards

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:47.206413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T08:07:34.112333Z digest=sha256:df8142a39271503d7fe0d4306276d3abfb3f15ef2f0cf63206dba5688031d9e4