Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T23:11:38.289892Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2608.09123.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T23:11:38.289892Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4e9b37a6-534e-43c0-a234-3ec5e4e92e13 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2889924e-d62e-4057-9eb2-c97356e32dfe · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a6f03e2-8876-42b4-b54c-3710f18f9f07 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning arXiv preprint arXiv:2511.12344 , year=
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e25f382d-7726-4dcd-8f98-3f52600eb963 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Qwen3 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72b2f941-3755-4bb3-a9a8-6261975bf125 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 08cc7842-2721-4c75-a21f-adc11d070f92 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b4fe6c1-b781-4a5f-a769-ec0739941a65 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Advances in Neural Information Processing Systems , volume=
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe2f8e82-8e4d-4241-9d0c-09b7c8812357 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d5fcef0-a15f-44e5-8771-6266d75b7602 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Rubric-Guided Self-Distillation: Post-Training Without Rubric Verifiers
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ae9b5fbc-19c2-4f56-a2a5-01ccf9ddf461 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bed4ef3-626f-49c7-a965-fa73d374b097 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84dfdcd7-c8fb-47c0-b765-55322c1a2eed · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28322207-e553-462d-91b6-485a2295f2c9 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Advances in Neural Information Processing Systems , volume=
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3bdef624-99cb-46d0-9025-8d16377d591f · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning HealthBench: Evaluating Large Language Models Towards Improved Human Health
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64de9da0-61a2-4473-8bb0-e97d61f24011 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83fa3cc0-f6ed-4374-a9c0-f610deb970b2 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Applied Sciences , volume=
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b5f20ae-92ad-4554-980c-951c20179e7e · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99bee403-91c7-4697-b25d-9935a7208b85 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Transactions of the Association for Computational Linguistics , volume=
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fc121aa3-0bed-4f75-af57-979f14a1cf01 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cef6ef6-5384-43ba-a09c-4d0c8b7a577e · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3de14e40-347c-4d9a-9978-5e867bbe0913 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Proceedings of the Twentieth European Conference on Computer Systems , pages=
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04377bda-206d-4a30-8fd3-fabdb4c3d6b6 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2b52015-9197-4f6a-948a-1c52fc66f927 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7258defc-00ed-4389-b722-b44d0ab08d68 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Qwen3-Coder-Next Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aee86933-812c-4885-8a8b-4da80ce8a7f7 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Advances in Neural Information Processing Systems , volume=
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83a8fd66-4828-47ab-8e32-3506cff0841e · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8e66481c-668f-48c6-b8d5-3e6499a81165 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning On the Creativity of Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c96373c-fbe9-4cd6-a317-02282b3da307 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 636e7c9e-4193-418b-8f5e-8bee23424b18 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Nature , volume=
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 91db8d45-85b1-4afa-a212-92576f8c04a9 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Nature medicine , volume=
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 241a1351-f8bc-4b7d-aa6a-a1e5294f9b31 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning SciQAG: A Framework for Auto-Generated Science Question Answering Dataset with Fine-grained Evaluation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58c720c6-929e-4461-896f-464740834da5 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Proceedings of the 34th ACM International Conference on Information and Knowledge Management , pages=
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6adbf349-02cd-4418-84e7-31ccbcb8747b · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Proceedings of the 58th annual meeting of the association for computational linguistics: system demonstrations , pages=
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c57b886f-9d57-441b-9441-3167ad8120cf · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume , pages=
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cb1e5e7e-dcdf-4ff3-86db-55733153b85e · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning LaMDA: Language Models for Dialog Applications
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdf34d31-bf96-4e2e-a42f-05c2132e88ae · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Advances in Neural Information Processing Systems , volume=
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ff5e8cbd-2f24-4256-ab36-2b320446ab19 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Reinforcement Learning via Self-Distillation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e05d968-6086-4871-b0ca-baa86e11f4bd · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37745c5f-89e8-4a9d-a9f0-a24d32e33a31 · outbound
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Findings of the Association for Computational Linguistics: ACL 2026 , pages=
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
No inbound Pith citation observations are available.