Pith. sign in

Paper Citation Record · LEDGER

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning

As of 7 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2604.21327.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.21327 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T22:55:19.471307Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact5
  • verified fuzzy3
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch13

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 993bcdb0-e451-4784-979e-d0e3b9e60613 · outbound

This paper cites AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:07.278352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:b6eae163ea4d6b67fa893132b70037ec5717b11f99befdedbf21b05c007672d5

Observation 853382d8-ffd6-40b4-895e-ebe07bbe1ab0 · outbound

This paper cites The Llama 3 Herd of Models.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning The Llama 3 Herd of Models

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T14:16:07.310219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:8e126c758e15766ab61ca01a2ac91d63aa20a33734185eae8ca24bce84a1bdbe

Observation 7e8e9d1e-c1e8-4ce8-96c1-20ca7ff6416e · outbound

This paper cites OpenAI o1 System Card.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning OpenAI o1 System Card

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T14:16:07.142595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:2095b6707a45bda370e1e4f37277a974eb4a4f5917d7cff6cce085808e1d20c5

Observation b004abd2-4277-4aac-99f7-ca41408c70fb · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T14:16:07.154064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:52a64b159f41e362b8fff961305d3e8ea33bddfa64c463e8aedf37331b917a9c

Observation 9a1177b7-7814-4e81-a5e1-4ebf023f189a · outbound

This paper cites ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:07.228056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:e5181961c1977b6e5199dbac0daac306f3c39d87a1965375f2de7eb448381c8c

Observation 7838304b-0556-4016-b531-a576edb2b8a6 · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning Maximizing Confidence Alone Improves Reasoning

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:07.165280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:c6b7e3832cbf86b98d456749d042a3a3726896d17a4bba711fc2b11033cf58c1

Observation 53c6ef3a-5f2d-43e1-8940-52451893aa05 · outbound

This paper cites Qwen2.5 Technical Report.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning Qwen2.5 Technical Report

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T14:16:07.195702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:9b62b035c3796b9bc6ca11878d03876dfc079b4bd6965e5172fa5bfb4bd1faaa

Observation 22c26f1a-5272-4615-b067-806161e7cc74 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T14:16:07.358998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:48aed9c62c47186559fe0311b9f355cc92d2362c43cdaca9b9dfc7b645ac8021

Observation 4ada04d9-6c8e-46a0-91d7-b097450ce33d · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T14:16:07.329408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:7828812ee12e80bee03b4dee163ffcc5f18284b9355611f73d9b24c667114899

Observation 3272d8e9-62e4-4bc8-95aa-6b7d80aaa11b · outbound

This paper cites Associated with the WaltonFuture GeoQA-8K-direct-synthesizing dataset release.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning Associated with the WaltonFuture GeoQA-8K-direct-synthesizing dataset release

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:07.183229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:b4e6ae2615f8009df807695d988aadcd90b1983858550d7e9ca1a48d54a9aa59

Observation ebf84b02-b72e-442b-848c-bd800b138f27 · outbound

This paper cites Dong Yan, Gaochen Wu, and Bowen Zhou.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning Dong Yan, Gaochen Wu, and Bowen Zhou

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:07.336008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:2348aacf1b9023bc7c81cab599c49384f57cee792aaf0802ab7952dfef9819ad

Observation a2c8f175-abc0-4a88-ac49-bc7f09d574ca · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:16:07.343427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:67909af8f7dbc829636245c7af21ac1bc3cd0a0356121295e1039808a07be24e

Observation df0842d8-0d3e-4837-bc14-c47570cad3ec · outbound

This paper cites Benchmarking Test-Time Adaptation against Distribution Shifts in Image Classification.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning Benchmarking Test-Time Adaptation against Distribution Shifts in Image Classification

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:07.318067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:63b5e1f2096403a2c2042f1a6a9525a2be14ce42f2ff2c6ade6f2bfa5c2a511e

Observation 1b1d1d6c-e1c3-4414-b792-8952e4fd56f8 · outbound

This paper cites Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:07.293158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:d48a16a2ab3ebca7c369358cb378a6416e4579d8587c30269e0ed4baa5f74694

Observation 40296b30-7cb6-48b9-a5ba-682c59888317 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T14:16:07.352061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:896aa2484e4925220390514e572f1955745a65df0fd419e9031a4ce7cd3e8248

Observation 98c45d61-939f-4e99-8b5e-db97e61b2730 · outbound

This paper cites Cumulative Reasoning with Large Language Models.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning Cumulative Reasoning with Large Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:43:16.267428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:d317f9946adc552fd4f03f359e650d9cd965b4d1ad54aa4b98f128c617206169

Observation 47ac53db-7c7a-400f-857d-031e7774a745 · outbound

This paper cites Learning to Reason without External Rewards.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning Learning to Reason without External Rewards

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T21:16:57.297836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:4536ea6cb0d23f17d28ca6be31aefae1d63f1e8cece690abfd3f6c5e01ae2b01

Observation 02984892-d7eb-4381-9052-b90a5366d86d · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning TTRL: Test-Time Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:07.213973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:55177658aef389a50f9e8a424dac0b9bf92570954136fd9d2ee5018e3bf76784

Observation 636bcd7c-dbdc-41d3-9723-ce41cf498bfe · outbound

This paper cites an unresolved cited work.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-23T13:58:02.718328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:aa3211bd8ae012f811dfe3fbd35657169a03bce016bc2f9e35444bba736362ba

Observation a352f429-65b4-483a-8dec-1fb65aad2099 · outbound

This paper cites This confirms that the key assump- tion behind BCS is not specific to MATH-500 but generalizes to other datasets.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning This confirms that the key assump- tion behind BCS is not specific to MATH-500 but generalizes to other datasets

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T13:58:02.714862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:209721bd48b3cbbeb71e46544133a9a5a6b5e55940955592c7552fbbd7a64b83

Observation cf37b11c-444c-4623-b886-3433621892fb · outbound

This paper cites To investigate the impact of the refinement size on model performance, we conduct a sensitivity analysis by scaling the original setting (1x) to 2x and 4x Table.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning To investigate the impact of the refinement size on model performance, we conduct a sensitivity analysis by scaling the original setting (1x) to 2x and 4x Table

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T13:58:02.707784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:4282879b544c14f945698f41bffe5b86d3d463d108747564743efeb0adcb030c

Observation 377fa298-9343-4b83-9f0b-a11375230cce · outbound

This paper cites Scaling up the refinement size yields highly marginal improvements across all benchmarks.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning Scaling up the refinement size yields highly marginal improvements across all benchmarks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T13:58:02.711142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:0823437695f0f5b37eed8e2172b0b74fc080873594d19b40de8b02a186833571

Pith citing papers

No inbound Pith citation observations are available.