Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-25T07:39:02.616294Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 15 inbound Pith citation observations for arXiv:2510.00915.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-25T07:39:02.616294Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:42:55.896502Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T17:09:58.372131Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 98a42c9f-8dd1-4fd8-86df-9b419ee7c54f · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Humans or llms as the judge? a study on judgement bias
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5b1f62e2-a05e-48e8-be83-1d5cfb58eeac · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Training Verifiers to Solve Math Word Problems
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6cf54ce2-4b73-4090-ace2-891ee0fa57c3 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers A Survey on LLM-as-a-Judge
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cfc6d225-1581-4626-bc4a-a22d52aad147 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Tsang, and Masashi Sugiyama
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ec6db028-df06-4282-b6c0-6ea20784e4b8 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Olympiadbench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 13751c81-e454-48d6-8873-9e1569bb6b39 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 672b0dbb-c5a9-4d03-815c-bee38f6723f6 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Pitfalls of rule- and model-based verifiers–a case study on mathematical reasoning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 67004685-a161-459d-aeb9-fdb090bdf9a4 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Math-verify: A robust mathematical expression evaluator for llm outputs
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fe4af45a-803b-4425-b277-1b2dbbe0ec68 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Aime 2024 (dataset card)
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5b15f4a9-442c-4e69-861f-d0618433d0d1 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 78e1f08f-5f86-4ebe-912e-db9b436c5afb · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers On the admissibility of horvitz-thompson estimator for estimating causal effects under network interference
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 393979c1-d642-4b18-a3de-f05aed788606 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dc68112d-0c0d-482a-955e-bdd1df293ef6 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fdc0212d-58bf-4c1f-94a1-ffe23d5ad8f8 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Provably end-to- end label-noise learning without anchor points
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0f8dab58-f029-481f-9742-466f818d888a · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0a53b119-b40e-463c-8796-818143a268c9 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Let’s verify step by step
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 436a6792-4ac2-497b-8b72-dbd9ee9552c7 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9770835f-3db2-4180-848a-6a87e0f73ce9 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Amc 2023 (dataset card)
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bf1f2286-cc07-4f85-8e17-eb0a1772070c · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Reinforcement learning with verifiable rewards: Grpo’s effective loss, dy- namics, and success amplification
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3911a067-e284-4931-a6a8-3ed17e9ea3d7 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Dhillon, Pradeep Ravikumar, and Ambuj Tewari
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7e1e6cca-c539-4fcc-a26d-0d2178f5d2de · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Aime 2025 (dataset card)
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d330f054-da8d-4ba9-81bc-7e1a3e718552 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Making deep neural networks robust to label noise: A loss correction approach
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 23225b11-a7e4-48b5-b6f7-b4010f1d6d83 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6ee03e51-64fc-4af9-8e3d-2303cf4c3e9e · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Optimization-based prompt injection attack to llm-as-a-judge
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ba3d881c-324c-4ce2-a3cf-d7f3fcc49825 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Judging the judges: A systematic study of position bias in llm-as-a-judge
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f61f0ddc-b7cf-42b9-960b-a0f3b08efaac · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Learning from noisy labels with deep neural networks: A survey.IEEE transactions on neural networks and learning systems, 34(11):8135–8153
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 803f7d05-7a98-4018-b59c-8439205fd533 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Sutton, David A
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b6917440-946e-479c-b0d9-ecf3803b79ec · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 766537cd-43be-420d-99c5-1dd1ffeea564 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Reinforcement learning with perturbed rewards
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 62f00431-e7b5-45eb-88e1-d987b8b61f4e · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Le, Ed H
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 659faf64-8df3-496c-95ca-e5846f62ef08 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Chi, Quoc V
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7b46c176-c529-496e-a393-add5908b0750 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 961c2790-f0b7-40dc-b5f0-5e5033e93272 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Simple statistical gradient-following algorithms for connectionist rein- forcement learning.Machine Learning, 8(3):229–256
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3e48e768-8302-413e-96bd-f89841ee7373 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 03019534-bcb7-4898-a55c-d0087e35b302 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Tree of thoughts: Deliberate problem solving with large language models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 64ea9ca2-05db-47ec-b342-b7b6cf48bc8d · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers One Token to Fool LLM-as-a-Judge
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d2c87dc9-fff0-4824-93a2-d41d09d8393a · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Le, and Ed H
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7efcb9d8-13b8-4993-a24b-ccb40a4b1340 · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e37b77f1-0056-4641-b1de-d252703fb3af · outbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers idx": 16
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 654e67cf-41d6-413b-9624-4a473739ed90 · inbound
Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7cd3a80d-c4cb-4415-91c8-c3456a6b27fe · inbound
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 867d0078-04a0-4ca3-a092-f0847c039b03 · inbound
Safe Bilevel Delegation (SBD): A Formal Framework for Runtime Delegation Safety in Multi-Agent Systems Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8e146cf6-da76-484f-8e82-84205622ea8f · inbound
Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0b992994-1517-4b56-86f4-b551dfc53705 · inbound
High-Dimensional Statistics: Reflections on Progress and Open Problems Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f24dfe78-ec82-400c-9273-e4294abc199b · inbound
High-Dimensional Statistics: Reflections on Progress and Open Problems Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 90abc5e3-fd0e-4dcf-86dd-f9bd22d5553c · inbound
On Training in Imagination Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3201df72-89be-48d6-a86c-e32e32d626d1 · inbound
On Training in Imagination Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 79184117-1bb2-420d-a45a-d9eba5009c89 · inbound
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 174ff37c-7b05-47ae-8741-77b7a89b6440 · inbound
Quantifying Empirical Compute-Supervision Tradeoffs in RLVR Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e8c54ab-3f8d-4d17-8172-952d8a433308 · inbound
GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 25ed3ebf-92e8-4346-ad5a-a39ad9ba6a7a · inbound
Reinforcement Learning for Computer-Use Agents with Autonomous Evaluation Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8f3545b0-d73a-4bde-8b0a-e2a7dab37ef9 · inbound
When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66b87979-c00d-42da-bec4-df343babbee9 · inbound
The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b80d79f-183c-4525-94a0-aabd8181c5d5 · inbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.