Pith. sign in

Paper Citation Record · LEDGER

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation

As of 16 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2506.15068.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15068 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:48:47.919007Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59633c9a-84b7-4b34-bf4d-4b5148322299 · outbound

This paper cites Large Language Models for Mathematical Reasoning: Progresses and Challenges.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Large Language Models for Mathematical Reasoning: Progresses and Challenges

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.792003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.792003Z digest=sha256:8f3a8cd3f2e1c871ff6f4cf19ff4c5fbc567677a407068e2ebdfcac1066f57e9

Observation 42880985-535d-434f-8c33-503a6db4456e · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.824662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.824662Z digest=sha256:c8f82df3ad040646b5817634da9a1184146c9c2208f293b016a7e7f99393a1e7

Observation 82c7b8bb-6cd8-4b10-9960-85dbb0006565 · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.829718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.829718Z digest=sha256:57a35ce5467ca30a60a303dd30b2167167611bd44a7b605312d2f0b56c64d06c

Observation 9b0f817d-33ce-49ea-a5df-f73e5b90ac41 · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.833859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.833859Z digest=sha256:afc83ff47bfbed8367990559ade8d8f71a2ca6ac319e7f212d73611a298f8e8e

Observation 7cfdd1aa-03dc-4f08-9ddd-6574288410bf · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.838048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.838048Z digest=sha256:2d8887178ecb84ec76cf5593e6a9c3936d7505a4358a1c890830e75828b987c7

Observation 7b1ddaaf-0d5a-433e-a3ee-d08404b52603 · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 6

Resolution
verified exact
doi, observed 2026-08-15T19:48:47.956081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:48:46.842925Z digest=sha256:8182409be1be147355a2ee5966ebaaa07907b4364ceba395d94af3e551b40150

Observation c5652fe0-37c2-4443-9162-5b0b1bb0ad3d · outbound

This paper cites Can Large Language Models Be an Alternative to Human Evaluations?.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Can Large Language Models Be an Alternative to Human Evaluations?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.903746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.903746Z digest=sha256:ec38e319f60c18290446dd5e5374e7d0aa97b85e79037a324a2e5a09ab18a474

Observation ab4f6210-4423-471c-a539-eaf97fb7ffdd · outbound

This paper cites A Closer Look into Automatic Evaluation Using Large Language Models.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation A Closer Look into Automatic Evaluation Using Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.963076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.963076Z digest=sha256:c3aa9728608dc838b2e44ec0af6ae2fb7f8408856d70f7f2e5abc2adf7aefc24

Observation b38cd49c-ff9f-41cb-9bb8-dd8aa9d367b9 · outbound

This paper cites Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.967870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.967870Z digest=sha256:2e636b807d240870c4a9d8bc663ed12c4d64fbff5a73343253e94272fcd15587

Observation 70d93744-92bd-46a9-a678-06b9312ab7dd · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.972385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.972385Z digest=sha256:20ea08138eb54cdac7cc6348b7c413ca4a099593c2ef0c598998d4aeb4778935

Observation 9fa16c0b-d5f1-4873-8484-7c74e090c983 · outbound

This paper cites Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.976457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.976457Z digest=sha256:c13023f863742c18cc3b93249d5449ffb73ecd117a31aea78c6384f1e1aee93a

Observation 7a0908a3-b4e9-4944-8c41-2c9f90b74268 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.981211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.981211Z digest=sha256:64b24fb1600e2055c191748f447a2611c740b842d66c566986b1d046a1900be6

Observation c711cc16-073e-493f-a9dc-5fd00a35fd2f · outbound

This paper cites SummEval: Re-evaluating Summarization Evaluation.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation SummEval: Re-evaluating Summarization Evaluation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.985186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.985186Z digest=sha256:99ad9bd46cee0b0e2c77f165880145c843a423fb235da4588a43b69014359a82

Observation f9f89c67-2d0e-487b-8ab7-d9944da35723 · outbound

This paper cites ELI5: Long Form Question Answering.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation ELI5: Long Form Question Answering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.993204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.993204Z digest=sha256:9a2e6f8d5106fd6efd9f81914d089b4eab9a2c5bb91b0ff8bc49ce708d99d4e3

Observation f87806cf-b748-4071-8717-7ddfc286289e · outbound

This paper cites A Survey on LLM-as-a-Judge.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation A Survey on LLM-as-a-Judge

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.109728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.109728Z digest=sha256:501144b4a3e4081ce573514a07d4daaafee3306903d074a748fdbd76fc4745ea

Observation 4a4a841e-ca3a-4ddc-aa54-1cafb9661be7 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.152322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.152322Z digest=sha256:a35936511ae1de73c114421010985007e56dc92e7c53d59b6144f28f4239eceb

Observation efa151a7-b8e1-4f0f-80bf-09e365e4def7 · outbound

This paper cites Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.156442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.156442Z digest=sha256:d4095c4bfb49d67580aca1b6d451a1aa229eb90b8e67c8d90e2037842b00850b

Observation 4c742731-11ea-453a-b364-66c2639d2694 · outbound

This paper cites Hurdles to Progress in Long-form Question Answering.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Hurdles to Progress in Long-form Question Answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.160503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.160503Z digest=sha256:22abd49f2a1db357206008dd789f7dff2aed1c3b0d87ec648347518ef4d9a18f

Observation 07198add-f6ef-4dbe-a58c-f159f398f067 · outbound

This paper cites No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.164830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.164830Z digest=sha256:dfc8e56c2b4198ba48ba0f288856ae905c75254a03bf1ed8a4284809255ca350

Observation d7fa89f0-82dc-4452-b104-3b6221a3f719 · outbound

This paper cites Kullback and R.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Kullback and R

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.169094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.169094Z digest=sha256:c8470ee635c7ccef4f4fdd9dd4f4143053dde6deace012637dd250b0fdbda46e

Observation b19010c5-7788-4fde-bea5-fb3b124d0b0f · outbound

This paper cites LongForm: Effective Instruction Tuning with Reverse Instructions.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation LongForm: Effective Instruction Tuning with Reverse Instructions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.173159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.173159Z digest=sha256:5db5963b2affe6a135d1eea78f70023ce706836011668929bdba1902f83f7173

Observation b1ada493-ff02-4dd9-a496-6bf5f803497c · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation RewardBench: Evaluating Reward Models for Language Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.177700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.177700Z digest=sha256:7356a393907a8d991a9e8030cba74a3d9b7f9406babfc4d1c61dd535659a11cc

Observation 911a17c8-de7d-42c1-a6ef-a86504782f47 · outbound

This paper cites CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.196869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.196869Z digest=sha256:51cf90e740ec23e588eb06d33d0082534eba99d66417ea7da50b29ed86c08edf

Observation bb7991a5-270e-4e7d-b946-c45ccf307c73 · outbound

This paper cites Deep Reinforcement Learning for Dialogue Generation.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Deep Reinforcement Learning for Dialogue Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.266034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.266034Z digest=sha256:3f54d5d951af6cdf8d2882359f28e31a0b58ba031d0217e1cc236515d29a6ba5

Observation ffc1a410-d56b-412f-9094-dce7e371e508 · outbound

This paper cites PEDANTS: Cheap but Effective and Interpretable Answer Equivalence.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation PEDANTS: Cheap but Effective and Interpretable Answer Equivalence

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.341857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.341857Z digest=sha256:30061f9beb7c11defb176f07e4ac4bc097c20f4ddade09c3e89db78c350d1659

Observation 49f107d5-ae54-4120-9047-51de611f1b45 · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:48:49.072381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:48:47.346440Z digest=sha256:135bb864028aee95c682ab99f17190c733a1142ea0a316361ba83ddfae05a416

Observation df51d059-bcb6-4a3b-bbcb-72dac589f17c · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.349856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.349856Z digest=sha256:2b9be08305ea83e91c91e31d3cc84c0653fa4b403d415678a1a7e8413834b780

Observation 8c56b692-c4c2-4d25-8589-c10ec768514b · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:48:49.060450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:48:47.354409Z digest=sha256:351ce8a29d352c215ac55b6cab1ae450d5f7537411cb5e48fc276eb921c1723e

Observation 6e17e919-6fd1-41e5-8794-8c9214c39b57 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.358973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.358973Z digest=sha256:00ba7410d84fd3cbc3ef0755a8d51dcf6dcddb576c4a45913c391742d63db612

Observation 4291197c-ead8-4f98-9239-b0b637b29654 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Understanding R1-Zero-Like Training: A Critical Perspective

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.363358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.363358Z digest=sha256:b79ef8fdeba3f8bb30e30623a5520bd57c6eab86639abd2e8fa33a34aec8a552

Observation 8833ad02-2bce-4a3f-86a6-d04de55c472b · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.368219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.368219Z digest=sha256:a32e26cf84426f27a42466852160659960b9a9151da2521b03cba40517ba3c55

Observation 7b1cb00b-6330-4a72-8ca3-e3c80af6e88a · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.426148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.426148Z digest=sha256:0f49a4a874896025acc264cd0cc67eccb0f4b9f7f9c3224c64191678439311c8

Observation caf6aad0-cb9e-486c-90a8-55ce572490db · outbound

This paper cites Qwen2.5 Technical Report.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Qwen2.5 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.445193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.445193Z digest=sha256:05796533cdf6ddfc02706a5ef9a59d74737cfab5b693e8d97d05144cb08c4b20

Observation e0d3f46b-16be-4087-a909-0c6928e5e594 · outbound

This paper cites Factually Consistent Summarization via Reinforcement Learning with Textual Entailment Feedback.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Factually Consistent Summarization via Reinforcement Learning with Textual Entailment Feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.490328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.490328Z digest=sha256:c2d3c6539a8a34d2b5bc1ec670d5b164787adc4068eb0139ed63f803bc75a4cd

Observation 2172f24f-836d-4973-bc8c-088284703ebc · outbound

This paper cites Enhancing Text Classification through LLM-Driven Active Learning and Human Annotation.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Enhancing Text Classification through LLM-Driven Active Learning and Human Annotation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.494498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.494498Z digest=sha256:2e84feb8c6d1820c2c463a33cbd48c26531722b116ccdc0260feb3a47db0efff

Observation 0ed25aac-2353-47cb-8753-c3843f9b9a99 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Proximal Policy Optimization Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.498566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.498566Z digest=sha256:a7bb6b0b07dfde0b6bfd12624b854abc1dd3a1b5747f07a5d240611df8296964

Observation d8c6387f-e736-4847-8856-3b8f8f9b70ef · outbound

This paper cites A Survey of Deep Reinforcement Learning in Video Games.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation A Survey of Deep Reinforcement Learning in Video Games

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.502875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.502875Z digest=sha256:536a6c29a1d4dabf7fbc9425b0715f05c458391375e3986b889c9df830314167

Observation 21355d6d-53d0-42d9-89ce-3933300a6fc7 · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:48:49.043155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:48:47.508657Z digest=sha256:c91fc89bc22e14421349f27e0a751b02d53affe552383877e4f74e0f2cfc40bc

Observation 3343d9bf-eafe-4e63-80b1-18cb54921874 · outbound

This paper cites Learning to summarize from human feedback.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Learning to summarize from human feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.511830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.511830Z digest=sha256:a49f566137e75a5ca4f37dbd77904536b94d0c337a50746cf647e2120be550f7

Observation 398115cb-a38d-4ff8-9663-c0d3013a26f7 · outbound

This paper cites Hashimoto.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Hashimoto

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:48:49.030890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:48:47.515624Z digest=sha256:aa3e341eb5706f1aaceaf8463024a793e554e71dab17238c5260d09371a77cd6

Observation 72d469f1-4e11-40b5-8359-e58a2bcfc064 · outbound

This paper cites Hashimoto.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Hashimoto

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.518693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.518693Z digest=sha256:75d78453cbd4a5f6ab8acff6fa736038c34234fe416b7053e5814b76d1a9e0dd

Observation 412274bc-b360-4114-9db9-a1b5c0282261 · outbound

This paper cites Smith, Daniel Khashabi, and Hannaneh Hajishirzi.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Smith, Daniel Khashabi, and Hannaneh Hajishirzi

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.522256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.522256Z digest=sha256:34fac22ee0ef01296c8914d2c039a14fe64bb938074631c8494cd64239762f21

Observation 11899969-0add-487d-bb90-380a48fbdb83 · outbound

This paper cites Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.589286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.589286Z digest=sha256:1974ef636a8be58d487178fc9bc7506b107688a68636056fca5012dd02147074

Observation 92417eca-0019-483a-90f9-b093826c8ad8 · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:48:48.945108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:48:47.646879Z digest=sha256:f67da9b5645f0176083e282adcb45d87ac201d681f839d1bb1215d4de951e832

Observation 290d8a8f-01fb-4548-89df-a767ace4b939 · outbound

This paper cites Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.650992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.650992Z digest=sha256:6060884d0fb6f9490d7ef1c5601f1ecf48b4d2eca3b9bcc26cb6e234b4a67c7f

Observation a98bc16a-ab2f-4290-95e6-7b81c94ec08d · outbound

This paper cites Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.655647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.655647Z digest=sha256:323afef6fa00ff35a800f3d8fe25605cc40ad6efb9764466435918eb75f913b1

Observation c51569b5-6d49-412f-baa6-70af2a9c25f0 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.659845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.659845Z digest=sha256:29cebbbbb396ad1fc09217494b45a5d7dde974421e963bb64b64883c1c17fcb9

Observation 5608c357-1103-4bb8-950f-b36f355de604 · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.664012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.664012Z digest=sha256:a85eac56240c8b7e3e1ad5a89227d0522963c9f278a93b863c4484e44ae4674f

Observation 9e215979-2aa4-4017-9edc-3cc08fc24453 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation BERTScore: Evaluating Text Generation with BERT

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.667853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.667853Z digest=sha256:259d8867f2b685bf39fd11b996e9129db3ac8b4fc4d50c80264133068e57f26a

Observation 9ea4d148-32eb-46c8-83ca-a04a4bf37f9e · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.671746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.671746Z digest=sha256:fda16b4704c62dff286d4cdd45e46170a1d10d9f7413d2515b37139e1b3dadad

Observation 35925866-a29d-4310-a845-2157a35a91f0 · outbound

This paper cites LIMA: Less Is More for Alignment.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation LIMA: Less Is More for Alignment

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.676062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.676062Z digest=sha256:18116e1ea497fe97d9b9aea400c8fc26402f3ac32a4eff132a71a4bc6a38218f

Observation 9d2223ba-ca32-4984-a71d-2faf3e8d0469 · outbound

This paper cites Teaching-Assistant-in-the-Loop: Improving Knowledge Distillation from Imperfect Teacher Models in Low-Budget Scenarios.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Teaching-Assistant-in-the-Loop: Improving Knowledge Distillation from Imperfect Teacher Models in Low-Budget Scenarios

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.680838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.680838Z digest=sha256:8b59b93e75e295eb6467efcec42182cb1a83da7dbd8dac3bb41e32055e0befb6

Observation 51ad07ec-509b-42f6-98b4-1090207865f2 · outbound

This paper cites MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:48:48.164263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:48:47.684812Z digest=sha256:3a0b8f2d4f4a190d56ab985ca0781b2dcad494257a59d26a940be5ea56e6d2b4

Observation 98ba6f26-631e-4f9b-8f52-c5861c9c830e · outbound

This paper cites an unresolved cited work.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.776043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.776043Z digest=sha256:ccf06b1b439e5bf896ea7e350aadf483a86f1b216ded7f08cb7ecc20a70ec93b

Observation a61909ed-061a-4b0e-a02b-c77979fd6e2f · outbound

This paper cites online" 'onlinestring :=.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation online" 'onlinestring :=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.869226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.869226Z digest=sha256:6c0d3f3e96d7cedf4cc8e3d693e7aec9eff0cef5194eefc06b2ea61c7e8f15dc

Observation 729d134a-ba3a-4338-a043-824819ecca29 · outbound

This paper cites write newline.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation write newline

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.919007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.919007Z digest=sha256:9201720c08014450ca2bb72a2044710ca038c9c112dbaab3efd082b37a3420c1

Pith citing papers

No inbound Pith citation observations are available.