Pith. sign in

Paper Citation Record · LEDGER

Lessons from Training Grounded LLMs with Verifiable Rewards

As of 7 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 2 inbound Pith citation observations for arXiv:2506.15522.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15522 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:59:49.564233Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T08:07:34.112333Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:09:47.204727Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 35225b90-78d7-4752-be6b-ec1e6cd2e9f6 · outbound

This paper cites I apologize, but I couldn't find an answer to your question in the search results.

Lessons from Training Grounded LLMs with Verifiable Rewards I apologize, but I couldn't find an answer to your question in the search results

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:59:50.547041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:59:49.564233Z digest=sha256:7d58e62cc5be379dc7bfbcd1328317f64b9b4ebec6f8fea90c10755face8c96d

Observation 0230525b-aa15-4555-b9e9-234d6bde6b96 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Lessons from Training Grounded LLMs with Verifiable Rewards DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.123573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.123573Z digest=sha256:c5b5ff4b6085cc882ac9bc4b260f2e123a9bdf0482fa0e2691498ab60e155847

Observation 45617367-ef31-4fb2-9d91-6a7148d34bca · outbound

This paper cites Training Language Models to Generate Text with Citations via Fine-grained Rewards.

Lessons from Training Grounded LLMs with Verifiable Rewards Training Language Models to Generate Text with Citations via Fine-grained Rewards

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.358493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.358493Z digest=sha256:675298c44fd63b52d1ce95a2e6d9e3dc64a93356479ac682de48d75d5af4b23c

Observation 3daa139d-305f-433a-ab13-407f4c4408fc · outbound

This paper cites RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement.

Lessons from Training Grounded LLMs with Verifiable Rewards RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.438690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.438690Z digest=sha256:4a2aa9901b3d70a134a0b96d143e603e13bb031515dcc4eaf342f5c9961b4e4c

Observation 300387b0-b5d7-485e-a35a-6b66106dc482 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Lessons from Training Grounded LLMs with Verifiable Rewards Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.887449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.887449Z digest=sha256:d1ffb31fbf98df2fe0c37a904f9d0af45bd091dd526c467feb645fa29bdce6fb

Observation a7343cb3-fb7d-42ca-b9e2-6bbd434ee48e · outbound

This paper cites Attribute First, then Generate: Locally-attributable Grounded Text Generation.

Lessons from Training Grounded LLMs with Verifiable Rewards Attribute First, then Generate: Locally-attributable Grounded Text Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.945371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.945371Z digest=sha256:4bf0a6f24aeb8d7fc7731f60b284c8f130c3ce6ed0823c365204c06d979b487b

Observation 38d883f4-49ef-47d8-9173-84b1fcecba0e · outbound

This paper cites Qwen3 Technical Report.

Lessons from Training Grounded LLMs with Verifiable Rewards Qwen3 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:49.122177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:49.122177Z digest=sha256:c652a7bacf0ee19c9c46cd8764f3505d81a4008f7462e65aa86fa83254809fee

Observation 9a973654-1c0c-499d-bdfe-82e746fde4a2 · outbound

This paper cites Ground Every Sentence: Improving Retrieval-Augmented LLMs with Interleaved Reference-Claim Generation.

Lessons from Training Grounded LLMs with Verifiable Rewards Ground Every Sentence: Improving Retrieval-Augmented LLMs with Interleaved Reference-Claim Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:49.203531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:49.203531Z digest=sha256:1fe2789ea01ca7329b8c13665c112f7590632a9d74409cfc9d164b237b271bf5

Observation bc6b8e4e-2b09-48c3-8f94-c45a11e00523 · outbound

This paper cites RECOMP: Improving Retrieval-Augmented LMs with Compression and Selective Augmentation.

Lessons from Training Grounded LLMs with Verifiable Rewards RECOMP: Improving Retrieval-Augmented LMs with Compression and Selective Augmentation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:49.286759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:49.286759Z digest=sha256:acc601e13a03ef1a855ee4943163ace46632883eef40848389fde04fa66201e7

Observation 18acf74a-00a1-4c3b-9ccc-2daf31e15d24 · outbound

This paper cites Effective Large Language Model Adaptation for Improved Grounding and Citation Generation.

Lessons from Training Grounded LLMs with Verifiable Rewards Effective Large Language Model Adaptation for Improved Grounding and Citation Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:49.340475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:49.340475Z digest=sha256:aba386352b9268a71bbaa39ac55f34353de1d4204444f7af1f6ce7eaae95bf37

Observation 125741e5-9e1a-47bc-8313-69b1c8f423f7 · outbound

This paper cites Making Retrieval-Augmented Language Models Robust to Irrelevant Context.

Lessons from Training Grounded LLMs with Verifiable Rewards Making Retrieval-Augmented Language Models Robust to Irrelevant Context

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:49.441945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:49.441945Z digest=sha256:c80c9fd5ab644c42101c32cd4a1255775c82872b8c6033459ef2e8e698068d46

Observation d956c89f-400c-4804-9a74-e5be51119b0c · outbound

This paper cites In Webber, B.; Cohn, T.; He, Y .; and Liu, Y ., eds.,Proceedings of the 2020 Confer- ence on Empirical Methods in Natural Language Processing (EMNLP), 6769–6781.

Lessons from Training Grounded LLMs with Verifiable Rewards In Webber, B.; Cohn, T.; He, Y .; and Liu, Y ., eds.,Proceedings of the 2020 Confer- ence on Empirical Methods in Natural Language Processing (EMNLP), 6769–6781

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:59:50.767047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:59:48.580786Z digest=sha256:fb7561699bed6e1ff54fb8de7853e59864f94755360511b3baea93a0aed55007

Observation ee134af0-4c9d-4123-91d5-14d8f7f6ab1d · outbound

This paper cites Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

Lessons from Training Grounded LLMs with Verifiable Rewards Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.657029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.657029Z digest=sha256:b759476c1e489eeea83d507f0bcbd40be81c70a3b23b2bd36523ddb188eba888

Observation 9c66798b-d9aa-4d47-8cd1-81052fcb18f4 · outbound

This paper cites Training language models to follow instructions with human feedback.

Lessons from Training Grounded LLMs with Verifiable Rewards Training language models to follow instructions with human feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.756238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.756238Z digest=sha256:f54a6c532a4ace0ef4402b76ca7d2acb8908473a5ff4e1ad196ab7e88483fe20

Observation 96fcc4ea-7249-48d0-9135-ca1ca25c81ce · outbound

This paper cites Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection.

Lessons from Training Grounded LLMs with Verifiable Rewards Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:47.915712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:47.915712Z digest=sha256:98b6505f8e9094e207d188b10095109959e57eba2e9c18f462a93893339acc18

Observation 9fe83f6a-6464-4523-bb30-8bd7adb02a88 · outbound

This paper cites The Llama 3 Herd of Models.

Lessons from Training Grounded LLMs with Verifiable Rewards The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.240829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.240829Z digest=sha256:12b721eb3f2f53902ff068a3848da1a520fc680d1b3a6af15380a2064633bc18

Observation fb49ac51-e732-439e-9968-a068c4997491 · outbound

This paper cites ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning.

Lessons from Training Grounded LLMs with Verifiable Rewards ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:47.976127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:47.976127Z digest=sha256:4463a19a4b3cfb0a53957a0daead4f5a92dbe5b13e5e31b505d170b9b0f218e6

Pith citing papers

Observation 49264500-90d5-4536-82ab-ac3512717904 · inbound

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards cites this paper.

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards Lessons from Training Grounded LLMs with Verifiable Rewards

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:26:28.297833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T14:24:48.666197Z digest=sha256:c76485ca9503dcab4d668b4138ff49ca17f6b06ce6cb3507364da27ed61e8f07

Observation dc08d080-03c2-41cf-88a2-737c511970a8 · inbound

Do LLM Attribution Metrics Transfer? Auditing Retrieval-Augmented Generation Evaluation Across Datasets and Constructs cites this paper.

Do LLM Attribution Metrics Transfer? Auditing Retrieval-Augmented Generation Evaluation Across Datasets and Constructs Lessons from Training Grounded LLMs with Verifiable Rewards

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:47.206413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T08:07:34.112333Z digest=sha256:69f45d26735e535f705f373262c9caf7681404a046bf35f9ce7453c8bf538058