Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

As of 6 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 15 inbound Pith citation observations for arXiv:2510.00915.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.00915 v4

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-25T07:39:02.616294Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:42:55.896502Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T17:09:58.372131Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact12
  • verified fuzzy23
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98a42c9f-8dd1-4fd8-86df-9b419ee7c54f · outbound

This paper cites Humans or llms as the judge? a study on judgement bias.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Humans or llms as the judge? a study on judgement bias

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.455528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:d6189fc10a45d1f2ebb4e250f70ea3217788f16ccc95c7172a1fa44c9b6ee6db

Observation 5b1f62e2-a05e-48e8-be83-1d5cfb58eeac · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Training Verifiers to Solve Math Word Problems

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:40:28.751883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:5c5fbc16b0f032564cd790ac5e204fec062b960c2ea5d86945321da54da9d0e9

Observation 6cf54ce2-4b73-4090-ace2-891ee0fa57c3 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers A Survey on LLM-as-a-Judge

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:40:28.726804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:9449f9279489e041ce8e134ac230c7b52d034aa5ee500350414c7932e7a43c04

Observation cfc6d225-1581-4626-bc4a-a22d52aad147 · outbound

This paper cites Tsang, and Masashi Sugiyama.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Tsang, and Masashi Sugiyama

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.410710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:dcce5a402085374419569d7fa63228d65b6993f85395c5fa20710c581a2c522a

Observation ec6db028-df06-4282-b6c0-6ea20784e4b8 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Olympiadbench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.431262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:e9f57378f87f5c9932735cd363bf27d6d41c3b714a2b795765c6c09be0faef44

Observation 13751c81-e454-48d6-8873-9e1569bb6b39 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-25T07:40:29.407612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:e9746603874b28431c6eb9d975cb28d88cc9827e374e6de9fe394f2f795568a7

Observation 672b0dbb-c5a9-4d03-815c-bee38f6723f6 · outbound

This paper cites Pitfalls of rule- and model-based verifiers–a case study on mathematical reasoning.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Pitfalls of rule- and model-based verifiers–a case study on mathematical reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:40:28.688777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:76b9a15249d5ccef8ee687c7b8f02e5d06266937e291d2c96b30e273b4592197

Observation 67004685-a161-459d-aeb9-fdb090bdf9a4 · outbound

This paper cites Math-verify: A robust mathematical expression evaluator for llm outputs.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Math-verify: A robust mathematical expression evaluator for llm outputs

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.417539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:c084814533b09babdcb09033de54a09e89bf69fdc499a75653dd01db2c744199

Observation fe4af45a-803b-4425-b277-1b2dbbe0ec68 · outbound

This paper cites Aime 2024 (dataset card).

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Aime 2024 (dataset card)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.427426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:f414040c6cb8f0c4c59d9b131569d880d69f3f3509bac77e4e6169cf011730a4

Observation 5b15f4a9-442c-4e69-861f-d0618433d0d1 · outbound

This paper cites Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.424226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:273c0df40910ed84cfa599c87d0877415e128c677106a01381878e514d9b06f3

Observation 78e1f08f-5f86-4ebe-912e-db9b436c5afb · outbound

This paper cites On the admissibility of horvitz-thompson estimator for estimating causal effects under network interference.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers On the admissibility of horvitz-thompson estimator for estimating causal effects under network interference

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:40:28.699666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:3263de2f9e7a5f3dd33cc5ad6b7778bef36c90476176d32b40142b13384f3dc2

Observation 393979c1-d642-4b18-a3de-f05aed788606 · outbound

This paper cites Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.499409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:cf35850825d9b7c9df1601d395da2838a6e5e8ed34927bba893c9b14654bee1b

Observation dc68112d-0c0d-482a-955e-bdd1df293ef6 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-25T07:40:29.503167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:0b71b409ed08ea7c36d55e59e63f9766ece8773d6907bb1afbe763de11cbb381

Observation fdc0212d-58bf-4c1f-94a1-ffe23d5ad8f8 · outbound

This paper cites Provably end-to- end label-noise learning without anchor points.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Provably end-to- end label-noise learning without anchor points

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.484272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:6f70be75063603b2f6221677c84c6aedfa1fac3aa643b4c1c9a45ca1489c1478

Observation 0f8dab58-f029-481f-9742-466f818d888a · outbound

This paper cites VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:40:28.733858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:ae28202ac112ad2a829c325a07a03a471101997aa6843607f38f9ff4e62a59c9

Observation 0a53b119-b40e-463c-8796-818143a268c9 · outbound

This paper cites Let’s verify step by step.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Let’s verify step by step

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.487892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:4d108404357f7d8397fef957ec5cb7f768ed658aa4a37e29ca60abe55c7ea2c0

Observation 436a6792-4ac2-497b-8b72-dbd9ee9552c7 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.491510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:8c25b7e96485eafef7319573c04302b13df9a7fd7d1833516c4dfc4bb3e9eb89

Observation 9770835f-3db2-4180-848a-6a87e0f73ce9 · outbound

This paper cites Amc 2023 (dataset card).

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Amc 2023 (dataset card)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.480701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:ecc9c0043d47763c7b0dfceeabf75a890e535fbaf0af67167b9482d82f7d2779

Observation bf1f2286-cc07-4f85-8e17-eb0a1772070c · outbound

This paper cites Reinforcement learning with verifiable rewards: Grpo’s effective loss, dy- namics, and success amplification.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Reinforcement learning with verifiable rewards: Grpo’s effective loss, dy- namics, and success amplification

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:40:28.705425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:6071e570e1d4852c2da807f77d896bbcf8c8d2c80ceb26301df5445488fd3eab

Observation 3911a067-e284-4931-a6a8-3ed17e9ea3d7 · outbound

This paper cites Dhillon, Pradeep Ravikumar, and Ambuj Tewari.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Dhillon, Pradeep Ravikumar, and Ambuj Tewari

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.466642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:2c6337f861c3f6479b32bd612cdee868e0260ef1feec2965ecaf47f539a89a8f

Observation 7e1e6cca-c539-4fcc-a26d-0d2178f5d2de · outbound

This paper cites Aime 2025 (dataset card).

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Aime 2025 (dataset card)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.469960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:006bc41365da61b850663ccfd8b5081977c91919ad3e245eb77e49e7a6cfb773

Observation d330f054-da8d-4ba9-81bc-7e1a3e718552 · outbound

This paper cites Making deep neural networks robust to label noise: A loss correction approach.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Making deep neural networks robust to label noise: A loss correction approach

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.473787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:b56bd05a01970092bf6de547fe0e4253d9d4f2da23ae65f1237a3c937746ecf0

Observation 23225b11-a7e4-48b5-b6f7-b4010f1d6d83 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:40:28.758312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:2dfd35c501d2e8d14f04ed12bedbd9d28d9ff945e155d65af3c6fea4ca4c86a5

Observation 6ee03e51-64fc-4af9-8e3d-2303cf4c3e9e · outbound

This paper cites Optimization-based prompt injection attack to llm-as-a-judge.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Optimization-based prompt injection attack to llm-as-a-judge

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.463272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:d2c891c564a3a0bac2a3d1afd1e53b91fca53e3bcba52857dd4200f64abb4355

Observation ba3d881c-324c-4ce2-a3cf-d7f3fcc49825 · outbound

This paper cites Judging the judges: A systematic study of position bias in llm-as-a-judge.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Judging the judges: A systematic study of position bias in llm-as-a-judge

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:40:28.716419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:f92b58cbd896388a7291194df2fc9d3318fa34db30ad64503f276bc6753546cd

Observation f61f0ddc-b7cf-42b9-960b-a0f3b08efaac · outbound

This paper cites Learning from noisy labels with deep neural networks: A survey.IEEE transactions on neural networks and learning systems, 34(11):8135–8153.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Learning from noisy labels with deep neural networks: A survey.IEEE transactions on neural networks and learning systems, 34(11):8135–8153

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.450869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:6ff1bda1bc0ecbea76db2ffc09731839344f503e98d3f2831d095ea7a4e5503e

Observation 803f7d05-7a98-4018-b59c-8439205fd533 · outbound

This paper cites Sutton, David A.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Sutton, David A

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.459332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:9845c25bb68fecd69f39895eef415a1e5b11202f853a05dba1b539acd28c7bb4

Observation b6917440-946e-479c-b0d9-ecf3803b79ec · outbound

This paper cites Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:40:28.682665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:ee43cff27bc4b609f64d6246acdf25e08624c496a1f4b044ca24885ed93c7c07

Observation 766537cd-43be-420d-99c5-1dd1ffeea564 · outbound

This paper cites Reinforcement learning with perturbed rewards.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Reinforcement learning with perturbed rewards

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.477459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:a5bc4504992c32187e4dc03457ee8b7623c1156a207a45286cfe910b225a72dd

Observation 62f00431-e7b5-45eb-88e1-d987b8b61f4e · outbound

This paper cites Le, Ed H.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Le, Ed H

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.446624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:0d536497d0f6f5719077e7d59c9d4537350f1e0684efd8c945ef63f23b0a16c4

Observation 659faf64-8df3-496c-95ca-e5846f62ef08 · outbound

This paper cites Chi, Quoc V.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Chi, Quoc V

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.443002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:045a5d8ed702ab65ee17e1c41f0a28057380eb53ea5163b1c1dcbaf7d2a8f7a8

Observation 7b46c176-c529-496e-a393-add5908b0750 · outbound

This paper cites Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:40:28.693870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:dbd4f04ae51eab23fc224c65bf07f4312b218eed67f47208f3824ce2ab98e634

Observation 961c2790-f0b7-40dc-b5f0-5e5033e93272 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist rein- forcement learning.Machine Learning, 8(3):229–256.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Simple statistical gradient-following algorithms for connectionist rein- forcement learning.Machine Learning, 8(3):229–256

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.435272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:c06624041322d6e1ae0b33428c423089cabb2185cd21ccfaee9d944be6b127c1

Observation 3e48e768-8302-413e-96bd-f89841ee7373 · outbound

This paper cites TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:40:28.721581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:1223152a2c58ca7d167a8b16afba059781881ceddf5e3d7a7fbec358ac3239e5

Observation 03019534-bcb7-4898-a55c-d0087e35b302 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Tree of thoughts: Deliberate problem solving with large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.439389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:5ee320873d5d8ec09a0bdb746221b3f233d30280915287278436a3b03a99c75c

Observation 64ea9ca2-05db-47ec-b342-b7b6cf48bc8d · outbound

This paper cites One Token to Fool LLM-as-a-Judge.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers One Token to Fool LLM-as-a-Judge

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-12T02:08:19.458599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:29d0e9e340d1320a0900c44cc632908b08bcf8e87699c236b0fecba447378b25

Observation d2c87dc9-fff0-4824-93a2-d41d09d8393a · outbound

This paper cites Le, and Ed H.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Le, and Ed H

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T07:40:29.495794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:86a814cb786ce0adadc04edebbb7f04130000c9fc84826d30baf506ba4413bd8

Observation 7efcb9d8-13b8-4993-a24b-ccb40a4b1340 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-25T07:40:29.420861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:e7ca17baf7ba00e276d8a99c403174f82b499fe4f2b2033995fc0e1fcf1c079c

Observation e37b77f1-0056-4641-b1de-d252703fb3af · outbound

This paper cites idx": 16.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers idx": 16

Reference 39

Resolution
malformed identifier
raw_fallback, observed 2026-05-25T07:40:29.414055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:ea2aee4074bd00e2184815805146a1b6eb41607f06bd075840fd664ab257358d

Pith citing papers

Observation 654e67cf-41d6-413b-9624-4a473739ed90 · inbound

Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing cites this paper.

Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-25T03:01:06.910934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T08:24:50.793194Z digest=sha256:acdd2997a45a7bd08508e2409fa24114c5abc2621e8bcc72b8546e24adaaac22

Observation 7cd3a80d-c4cb-4415-91c8-c3456a6b27fe · inbound

VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction cites this paper.

VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:26:41.698773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:26:35.767179Z digest=sha256:095c56d76edb8f9d794081fc5ea5e3f8bbe87cc344fe8dcba9ab7cc8a4b31536

Observation 867d0078-04a0-4ca3-a092-f0847c039b03 · inbound

Safe Bilevel Delegation (SBD): A Formal Framework for Runtime Delegation Safety in Multi-Agent Systems cites this paper.

Safe Bilevel Delegation (SBD): A Formal Framework for Runtime Delegation Safety in Multi-Agent Systems Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:01:06.910934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T08:03:21.322567Z digest=sha256:341bb6db9c33f7c00849464853e4e99f23370aa498eab4bb210823b2295d6708

Observation 8e146cf6-da76-484f-8e82-84205622ea8f · inbound

Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR cites this paper.

Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:01:06.910934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T18:52:52.969408Z digest=sha256:e207e8ce43b0658721e3f12871d229e1485036599769e364fef5995ac66c6391

Observation 0b992994-1517-4b56-86f4-b551dfc53705 · inbound

High-Dimensional Statistics: Reflections on Progress and Open Problems cites this paper.

High-Dimensional Statistics: Reflections on Progress and Open Problems Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:01:06.910934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T15:35:08.202464Z digest=sha256:c10b624bbee97c451963ea6fa43153b33b8e9c1f2a4d38112379e663ae0ed044

Observation f24dfe78-ec82-400c-9273-e4294abc199b · inbound

High-Dimensional Statistics: Reflections on Progress and Open Problems cites this paper.

High-Dimensional Statistics: Reflections on Progress and Open Problems Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:15:46.696413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T23:25:55.375541Z digest=sha256:14924c5fb9a37dbac8e0931f7f5f95fdd43c4b11fe52da485e9f609148b2b280

Observation 90abc5e3-fd0e-4dcf-86dd-f9bd22d5553c · inbound

On Training in Imagination cites this paper.

On Training in Imagination Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-25T03:01:06.910934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T01:00:03.971064Z digest=sha256:3aa8a1d1bd9821aa70e84caaf0ad30d4c4dac0b1febde14830e30e290d1304d7

Observation 3201df72-89be-48d6-a86c-e32e32d626d1 · inbound

On Training in Imagination cites this paper.

On Training in Imagination Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:01:06.910934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:35:12.364580Z digest=sha256:3b55c83243d1e645b655e9658f4f95a87b8446942c6fcb6e48ca14050a2cd629

Observation 79184117-1bb2-420d-a45a-d9eba5009c89 · inbound

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR cites this paper.

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:01:06.910934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T05:02:35.271960Z digest=sha256:0386b1cb0241adfe659a2b59dc930c8864873cdd3316300d42302110096b563b

Observation 174ff37c-7b05-47ae-8741-77b7a89b6440 · inbound

Quantifying Empirical Compute-Supervision Tradeoffs in RLVR cites this paper.

Quantifying Empirical Compute-Supervision Tradeoffs in RLVR Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T11:54:38.515532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T11:46:34.688973Z digest=sha256:576cfce131e9e99069c1e7c27c75788e82d7b16decd48bb9e04f5b43fd85d462

Observation 9e8c54ab-3f8d-4d17-8172-952d8a433308 · inbound

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling cites this paper.

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T06:06:40.805627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T07:45:43.320339Z digest=sha256:7cc4e596d155c437f28a155c475764ad495a26630ce35bda2aac7cf260e2bdb4

Observation 25ed3ebf-92e8-4346-ad5a-a39ad9ba6a7a · inbound

Reinforcement Learning for Computer-Use Agents with Autonomous Evaluation cites this paper.

Reinforcement Learning for Computer-Use Agents with Autonomous Evaluation Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:09:58.373417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T23:58:48.558959Z digest=sha256:f1fd2c091eb6221848ffaeae05d28bd525b2a8507598b24bea90cae6896e3cee

Observation 8f3545b0-d73a-4bde-8b0a-e2a7dab37ef9 · inbound

When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR cites this paper.

When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T07:36:47.336670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:36:47.336670Z digest=sha256:f2c4d1d8cf6f1f188c3a35abe96d69efc7979ba1c09169614e586bcb2d8e15bf

Observation 66b87979-c00d-42da-bec4-df343babbee9 · inbound

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy cites this paper.

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T06:30:16.612345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:30:16.612345Z digest=sha256:fd9d2701f24796d73c78b5d9a42502729210f2565c2ef0619210c12a63146129

Observation 1b80d79f-183c-4525-94a0-aabd8181c5d5 · inbound

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets cites this paper.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:55.896502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:55.896502Z digest=sha256:a1ffbbb14717a5dae9ad873b9b3be8a38e45d9901745c1ec929900617a4fcf11