Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:59:49.564233Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 2 inbound Pith citation observations for arXiv:2506.15522.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:59:49.564233Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-26T08:07:34.112333Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T11:09:47.204727Z
17 of 17 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 35225b90-78d7-4752-be6b-ec1e6cd2e9f6 · outbound
Lessons from Training Grounded LLMs with Verifiable Rewards I apologize, but I couldn't find an answer to your question in the search results
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0230525b-aa15-4555-b9e9-234d6bde6b96 · outbound
Lessons from Training Grounded LLMs with Verifiable Rewards DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45617367-ef31-4fb2-9d91-6a7148d34bca · outbound
Lessons from Training Grounded LLMs with Verifiable Rewards Training Language Models to Generate Text with Citations via Fine-grained Rewards
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3daa139d-305f-433a-ab13-407f4c4408fc · outbound
Lessons from Training Grounded LLMs with Verifiable Rewards RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 300387b0-b5d7-485e-a35a-6b66106dc482 · outbound
Lessons from Training Grounded LLMs with Verifiable Rewards Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7343cb3-fb7d-42ca-b9e2-6bbd434ee48e · outbound
Lessons from Training Grounded LLMs with Verifiable Rewards Attribute First, then Generate: Locally-attributable Grounded Text Generation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38d883f4-49ef-47d8-9173-84b1fcecba0e · outbound
Lessons from Training Grounded LLMs with Verifiable Rewards Qwen3 Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a973654-1c0c-499d-bdfe-82e746fde4a2 · outbound
Lessons from Training Grounded LLMs with Verifiable Rewards Ground Every Sentence: Improving Retrieval-Augmented LLMs with Interleaved Reference-Claim Generation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc6b8e4e-2b09-48c3-8f94-c45a11e00523 · outbound
Lessons from Training Grounded LLMs with Verifiable Rewards RECOMP: Improving Retrieval-Augmented LMs with Compression and Selective Augmentation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18acf74a-00a1-4c3b-9ccc-2daf31e15d24 · outbound
Lessons from Training Grounded LLMs with Verifiable Rewards Effective Large Language Model Adaptation for Improved Grounding and Citation Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 125741e5-9e1a-47bc-8313-69b1c8f423f7 · outbound
Lessons from Training Grounded LLMs with Verifiable Rewards Making Retrieval-Augmented Language Models Robust to Irrelevant Context
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d956c89f-400c-4804-9a74-e5be51119b0c · outbound
Lessons from Training Grounded LLMs with Verifiable Rewards In Webber, B.; Cohn, T.; He, Y .; and Liu, Y ., eds.,Proceedings of the 2020 Confer- ence on Empirical Methods in Natural Language Processing (EMNLP), 6769–6781
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee134af0-4c9d-4123-91d5-14d8f7f6ab1d · outbound
Lessons from Training Grounded LLMs with Verifiable Rewards Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c66798b-d9aa-4d47-8cd1-81052fcb18f4 · outbound
Lessons from Training Grounded LLMs with Verifiable Rewards Training language models to follow instructions with human feedback
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96fcc4ea-7249-48d0-9135-ca1ca25c81ce · outbound
Lessons from Training Grounded LLMs with Verifiable Rewards Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fe83f6a-6464-4523-bb30-8bd7adb02a88 · outbound
Lessons from Training Grounded LLMs with Verifiable Rewards The Llama 3 Herd of Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb49ac51-e732-439e-9968-a068c4997491 · outbound
Lessons from Training Grounded LLMs with Verifiable Rewards ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49264500-90d5-4536-82ab-ac3512717904 · inbound
Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards Lessons from Training Grounded LLMs with Verifiable Rewards
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc08d080-03c2-41cf-88a2-737c511970a8 · inbound
Do LLM Attribution Metrics Transfer? Auditing Retrieval-Augmented Generation Evaluation Across Datasets and Constructs Lessons from Training Grounded LLMs with Verifiable Rewards
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.