Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2606.09883.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9c019f15-2c7f-49e2-a140-3517abf27c08 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 41fcf300-4b88-4686-8e7e-c65e293bb56a · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17c627f7-10fa-4bb0-b787-a1c00e68004c · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06adb225-299a-4987-a378-8ada585435be · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a41f56a-ad4e-4110-b55d-7f925c6ec038 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64f1be4b-8902-4965-be48-a6e93e12ac9a · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Output only the requested text blocks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30a195a5-ebf0-4583-a9b9-190a4de8fdc0 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79fdd6e5-a81e-409e-9292-21cd71056276 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Do not refer to other subproblems, previous results, or hidden context
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4783e974-6592-4942-8b47-9bb8f403e986 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition This restatement can be natural prose or an explicit ‘Given:‘ clause
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05951660-e96c-4b33-bec4-9086cd0230be · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition show that
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86f6260b-b9a0-4618-89f4-f957510b3350 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition is the final answer correct?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1996ca7c-318a-4923-88a9-21a0a6ff5c53 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Do not invent abstract variables like S, T, or R unless they are defined inside the same subproblem and genuinely useful
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcedcf2f-afea-4ddd-bab1-97d7906b2cee · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Do not include explanations, equations, or multiple sentences in ‘answer‘
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46b6f42a-9a73-4555-aba5-2fa3d34bb276 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Never leave any field blank
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e87a1662-4257-4c61-8ec8-89a3afbb1d1a · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 747466d0-2284-4e30-8f28-7b4dcf758868 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Avoid repeating the same shell sentence with only numbers changed
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ae2eff4-3ca7-49a1-b038-81fe39ccc5f5 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition If the problem feels simple, split it into smaller concrete computations anyway
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f06dee5-9949-4f44-b25a-2d545a59de61 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4782c24c-1bb2-467d-b515-1a72f09880f1 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Each subproblem should ask for one concrete intermediate quantity, relation, or check
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8eca67b-4182-4449-992f-7535029fd602 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69ca7f9c-7778-42d5-8b9e-dfc93913b010 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a47a98e-9c73-403a-b51b-7c91a889ea33 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Do not leave ‘solution ‘, ‘answer‘, or ‘verification‘ blank
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e50a96-577e-4979-ac1c-ffac51372424 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Do not turn a subproblem into a mini-lecture or a long multi-part derivation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf258155-77cd-4141-aeb7-13daec0b6d1c · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Use 7 or 8 only when the original problem clearly needs them
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5034a7e8-d837-4849-a3b1-46f59071f37e · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Never write phrases like ‘from the previous step‘, ‘from Subproblem 3‘, ‘using the result above‘, or ‘same as before‘
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19370d56-c767-4887-80ec-b72c12a6607d · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition The final subproblem must always contain a non-empty ‘solution‘ and a non-empty ‘answer‘
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73faa848-2e28-4bf3-ade7-b4c21c980526 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition If the original answer is an expression, give only that expression
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9af39f7-598c-4ded-a048-c504c3822add · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition If the last step would only check correctness, merge it into the previous computational step and keep the final answer there
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb0cec48-91c8-4f63-87ee-4be6beae670c · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Do not write explanatory prefixes such as ‘Therefore‘, ‘So‘, ‘The answer is ‘, ‘check‘, ‘because‘, or a full sentence
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d71de62-e3b2-4068-8495-df8471fcb826 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a589746-4b9d-4003-8223-67cd994012ec · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b3f24fe-3d5a-4d07-8814-eb55a702cb0b · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9be0412e-d611-4885-b3b4-fb59c8978883 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d55759f-67c0-4e2a-b659-afc03e68f1b8 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e8cf8cf-78d0-46c0-870a-db8fe712ad51 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 711d9049-3b33-4cd5-8bba-73bab9ee1a61 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9017a485-850b-415f-8045-70c0e3263f34 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5975ba5e-8dce-41f6-9808-27977a950522 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5172a46a-5c14-4984-b1c0-f2d554865378 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 790696ad-21be-42dd-bb3d-10d92313f679 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7c79992-5738-494a-bb16-aa3e581e2980 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d076cd4-016d-457b-890f-2425aff52f11 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96f61686-b245-40d7-9e0c-581bfb9ea0bb · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8414f0d0-907b-4e98-ba8b-56eab90d50db · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c15dfecf-6b25-46aa-8325-26554c3fe840 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition results": [ {
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92b41ef1-4290-47ed-b83d-6d0fcc6c5856 · outbound
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition subproblem accuracy goes up, therefore root accuracy goes up
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.