Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:48:47.919007Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2506.15068.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:48:47.919007Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 59633c9a-84b7-4b34-bf4d-4b5148322299 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Large Language Models for Mathematical Reasoning: Progresses and Challenges
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42880985-535d-434f-8c33-503a6db4456e · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82c7b8bb-6cd8-4b10-9960-85dbb0006565 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b0f817d-33ce-49ea-a5df-f73e5b90ac41 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cfdd1aa-03dc-4f08-9ddd-6574288410bf · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b1ddaaf-0d5a-433e-a3ee-d08404b52603 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c5652fe0-37c2-4443-9162-5b0b1bb0ad3d · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Can Large Language Models Be an Alternative to Human Evaluations?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab4f6210-4423-471c-a539-eaf97fb7ffdd · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation A Closer Look into Automatic Evaluation Using Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b38cd49c-ff9f-41cb-9bb8-dd8aa9d367b9 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70d93744-92bd-46a9-a678-06b9312ab7dd · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fa16c0b-d5f1-4873-8484-7c74e090c983 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a0908a3-b4e9-4944-8c41-2c9f90b74268 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c711cc16-073e-493f-a9dc-5fd00a35fd2f · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation SummEval: Re-evaluating Summarization Evaluation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9f89c67-2d0e-487b-8ab7-d9944da35723 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation ELI5: Long Form Question Answering
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f87806cf-b748-4071-8717-7ddfc286289e · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation A Survey on LLM-as-a-Judge
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a4a841e-ca3a-4ddc-aa54-1cafb9661be7 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efa151a7-b8e1-4f0f-80bf-09e365e4def7 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c742731-11ea-453a-b364-66c2639d2694 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Hurdles to Progress in Long-form Question Answering
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07198add-f6ef-4dbe-a58c-f159f398f067 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7fa89f0-82dc-4452-b104-3b6221a3f719 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Kullback and R
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b19010c5-7788-4fde-bea5-fb3b124d0b0f · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation LongForm: Effective Instruction Tuning with Reverse Instructions
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1ada493-ff02-4dd9-a496-6bf5f803497c · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation RewardBench: Evaluating Reward Models for Language Modeling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 911a17c8-de7d-42c1-a6ef-a86504782f47 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb7991a5-270e-4e7d-b946-c45ccf307c73 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Deep Reinforcement Learning for Dialogue Generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffc1a410-d56b-412f-9094-dce7e371e508 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation PEDANTS: Cheap but Effective and Interpretable Answer Equivalence
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49f107d5-ae54-4120-9047-51de611f1b45 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation df51d059-bcb6-4a3b-bbcb-72dac589f17c · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c56b692-c4c2-4d25-8589-c10ec768514b · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6e17e919-6fd1-41e5-8794-8c9214c39b57 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4291197c-ead8-4f98-9239-b0b637b29654 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Understanding R1-Zero-Like Training: A Critical Perspective
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8833ad02-2bce-4a3f-86a6-d04de55c472b · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b1cb00b-6330-4a72-8ca3-e3c80af6e88a · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caf6aad0-cb9e-486c-90a8-55ce572490db · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Qwen2.5 Technical Report
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0d3f46b-16be-4087-a909-0c6928e5e594 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Factually Consistent Summarization via Reinforcement Learning with Textual Entailment Feedback
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2172f24f-836d-4973-bc8c-088284703ebc · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Enhancing Text Classification through LLM-Driven Active Learning and Human Annotation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ed25aac-2353-47cb-8753-c3843f9b9a99 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Proximal Policy Optimization Algorithms
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8c6387f-e736-4847-8856-3b8f8f9b70ef · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation A Survey of Deep Reinforcement Learning in Video Games
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21355d6d-53d0-42d9-89ce-3933300a6fc7 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3343d9bf-eafe-4e63-80b1-18cb54921874 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Learning to summarize from human feedback
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 398115cb-a38d-4ff8-9663-c0d3013a26f7 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Hashimoto
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 72d469f1-4e11-40b5-8359-e58a2bcfc064 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Hashimoto
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 412274bc-b360-4114-9db9-a1b5c0282261 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Smith, Daniel Khashabi, and Hannaneh Hajishirzi
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11899969-0add-487d-bb90-380a48fbdb83 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92417eca-0019-483a-90f9-b093826c8ad8 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 290d8a8f-01fb-4548-89df-a767ace4b939 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a98bc16a-ab2f-4290-95e6-7b81c94ec08d · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c51569b5-6d49-412f-baa6-70af2a9c25f0 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5608c357-1103-4bb8-950f-b36f355de604 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e215979-2aa4-4017-9edc-3cc08fc24453 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation BERTScore: Evaluating Text Generation with BERT
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ea4d148-32eb-46c8-83ca-a04a4bf37f9e · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35925866-a29d-4310-a845-2157a35a91f0 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation LIMA: Less Is More for Alignment
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d2223ba-ca32-4984-a71d-2faf3e8d0469 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Teaching-Assistant-in-the-Loop: Improving Knowledge Distillation from Imperfect Teacher Models in Low-Budget Scenarios
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51ad07ec-509b-42f6-98b4-1090207865f2 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 98ba6f26-631e-4f9b-8f52-c5861c9c830e · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a61909ed-061a-4b0e-a02b-c77979fd6e2f · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation online" 'onlinestring :=
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 729d134a-ba3a-4338-a043-824819ecca29 · outbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation write newline
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.