Pith. sign in

Paper Citation Record · LEDGER

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition

As of 5 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2606.09883.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.09883 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved43
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9c019f15-2c7f-49e2-a140-3517abf27c08 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T07:01:44.350935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:96dfc43b3f6190a31735119fbf4f2ab8c41cb4521c73f9e0d7612294c5f20104

Observation 41fcf300-4b88-4686-8e7e-c65e293bb56a · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:63073e2824238502412bb36edecdfdf21c9b0375515f3a9b4ddab320e9134044

Observation 17c627f7-10fa-4bb0-b787-a1c00e68004c · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 3

Resolution
parse uncertain
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:188c7689d2bb2163846902c87e9919ea05451404009457b55c6c60b177975fa6

Observation 06adb225-299a-4987-a378-8ada585435be · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:32f7063c8e1d4de4687743e8d7d1c710aa0bdf9d6516a4a8c14c5bbef888a230

Observation 5a41f56a-ad4e-4110-b55d-7f925c6ec038 · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:232f9e4eec3e3251827f5606d4d012580fd34e0d7836689f21fe66b0c6d085e1

Observation 64f1be4b-8902-4965-be48-a6e93e12ac9a · outbound

This paper cites Output only the requested text blocks.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Output only the requested text blocks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:74de7caac74ae1bacf4431569c4a716b77f6433c8be588810b09e62b76b5f28a

Observation 30a195a5-ebf0-4583-a9b9-190a4de8fdc0 · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:76f909dd5608dc48974de12bf7d48e73127ee6caace267f0be7ed9fb365d258f

Observation 79fdd6e5-a81e-409e-9292-21cd71056276 · outbound

This paper cites Do not refer to other subproblems, previous results, or hidden context.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Do not refer to other subproblems, previous results, or hidden context

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:761e456ac1740f5ba54cfccf9e6ab0e5e10f6bf8d13fd3f8ea94339b6bb68e56

Observation 4783e974-6592-4942-8b47-9bb8f403e986 · outbound

This paper cites This restatement can be natural prose or an explicit ‘Given:‘ clause.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition This restatement can be natural prose or an explicit ‘Given:‘ clause

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:dbba220e7632db5630a776c101eb3aa5cc1dfb58a5849339a1d63f8627b4a1b9

Observation 05951660-e96c-4b33-bec4-9086cd0230be · outbound

This paper cites show that.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition show that

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:aae50f7df712de0a4b94d605cf606826773658cc05776b65caf74f4e03ef2a93

Observation 86f6260b-b9a0-4618-89f4-f957510b3350 · outbound

This paper cites is the final answer correct?.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition is the final answer correct?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:f9e1880f112cecc05ce9a89c0a4e1836ee4e9cbc37290797912709895820f1f0

Observation 1996ca7c-318a-4923-88a9-21a0a6ff5c53 · outbound

This paper cites Do not invent abstract variables like S, T, or R unless they are defined inside the same subproblem and genuinely useful.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Do not invent abstract variables like S, T, or R unless they are defined inside the same subproblem and genuinely useful

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:3fdcb44b8fd2297396c87d56436863226fd07b62ddb4d1543ba184d386103eba

Observation dcedcf2f-afea-4ddd-bab1-97d7906b2cee · outbound

This paper cites Do not include explanations, equations, or multiple sentences in ‘answer‘.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Do not include explanations, equations, or multiple sentences in ‘answer‘

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:08967592ddf850951228b53e1a4ad04adc0cc8ec059f910b8e92ef50435cab8b

Observation 46b6f42a-9a73-4555-aba5-2fa3d34bb276 · outbound

This paper cites Never leave any field blank.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Never leave any field blank

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:2b5e7816e7c05c159f45369a2d2cf7f1181e53420b45357750f94e0aecec1d4c

Observation e87a1662-4257-4c61-8ec8-89a3afbb1d1a · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:52e5e02854d0c82c8a2f9d09bdbb4c7799fc4ab8931233c3be6049c5e2ab3745

Observation 747466d0-2284-4e30-8f28-7b4dcf758868 · outbound

This paper cites Avoid repeating the same shell sentence with only numbers changed.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Avoid repeating the same shell sentence with only numbers changed

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:c5553c632836612285b4733e564f2a542f35bb3b4ac93c1a0461efea2e9a170f

Observation 3ae2eff4-3ca7-49a1-b038-81fe39ccc5f5 · outbound

This paper cites If the problem feels simple, split it into smaller concrete computations anyway.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition If the problem feels simple, split it into smaller concrete computations anyway

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:4813ca9061c393375d568e885ed6e76a1004a26ab12790b410546b2ee597a9aa

Observation 6f06dee5-9949-4f44-b25a-2d545a59de61 · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:1d182439ed7c97095c34185085e043c51c9c31191a7fa991171a98019109d2bb

Observation 4782c24c-1bb2-467d-b515-1a72f09880f1 · outbound

This paper cites Each subproblem should ask for one concrete intermediate quantity, relation, or check.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Each subproblem should ask for one concrete intermediate quantity, relation, or check

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:e1a993bd3b056fb9e6cec901f549749aa43379e611f0c3810a2d4db856208aad

Observation d8eca67b-4182-4449-992f-7535029fd602 · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:267e961725e906e79214d40611f773cee7181412046b6caf22549be5d07880aa

Observation 69ca7f9c-7778-42d5-8b9e-dfc93913b010 · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:ddfea6206893845654369249994e4fccf47dc2a0004948e9a7a37d754c7df450

Observation 9a47a98e-9c73-403a-b51b-7c91a889ea33 · outbound

This paper cites Do not leave ‘solution ‘, ‘answer‘, or ‘verification‘ blank.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Do not leave ‘solution ‘, ‘answer‘, or ‘verification‘ blank

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:c0c013f02bec0cf67b590a521c22b1f525e26991b2773795d77f40c9635697c6

Observation c3e50a96-577e-4979-ac1c-ffac51372424 · outbound

This paper cites Do not turn a subproblem into a mini-lecture or a long multi-part derivation.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Do not turn a subproblem into a mini-lecture or a long multi-part derivation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:9f10908bbea49f8a7a43293408c7db26453696b3cb9884b778c6078695dc7321

Observation bf258155-77cd-4141-aeb7-13daec0b6d1c · outbound

This paper cites Use 7 or 8 only when the original problem clearly needs them.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Use 7 or 8 only when the original problem clearly needs them

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:f05d9936066242da384627239e21044393422195a3ae2b2f8140f68bf31f1b66

Observation 5034a7e8-d837-4849-a3b1-46f59071f37e · outbound

This paper cites Never write phrases like ‘from the previous step‘, ‘from Subproblem 3‘, ‘using the result above‘, or ‘same as before‘.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Never write phrases like ‘from the previous step‘, ‘from Subproblem 3‘, ‘using the result above‘, or ‘same as before‘

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:7e6cdd05eaccbe9317583ee8f6ff5070eabe3c00c223e0fecfb608dc6c4c1f14

Observation 19370d56-c767-4887-80ec-b72c12a6607d · outbound

This paper cites The final subproblem must always contain a non-empty ‘solution‘ and a non-empty ‘answer‘.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition The final subproblem must always contain a non-empty ‘solution‘ and a non-empty ‘answer‘

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:37be18ba7fd6df73ac4e265601e726859b3098d99f9c88b74c1c71902e9abd37

Observation 73faa848-2e28-4bf3-ade7-b4c21c980526 · outbound

This paper cites If the original answer is an expression, give only that expression.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition If the original answer is an expression, give only that expression

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:4bb89e328fed2bfda4bac9e02588492ad1a01a7f1b85304e6d6da71a739082ee

Observation f9af39f7-598c-4ded-a048-c504c3822add · outbound

This paper cites If the last step would only check correctness, merge it into the previous computational step and keep the final answer there.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition If the last step would only check correctness, merge it into the previous computational step and keep the final answer there

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:6b5aa8eb6a1411fc5a43683690c21a8b5eca79b06c2f6586f2b6eb7ea6bdccbb

Observation cb0cec48-91c8-4f63-87ee-4be6beae670c · outbound

This paper cites Do not write explanatory prefixes such as ‘Therefore‘, ‘So‘, ‘The answer is ‘, ‘check‘, ‘because‘, or a full sentence.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Do not write explanatory prefixes such as ‘Therefore‘, ‘So‘, ‘The answer is ‘, ‘check‘, ‘because‘, or a full sentence

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:dfca7a327507b61e8d344de7a9a1536eec6bf617fdb74bf06f5e906d2a518d49

Observation 2d71de62-e3b2-4068-8495-df8471fcb826 · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:2346ef6f8c1f291d5b07983e95b818d9766ab125466981cee9c195fc688c36bc

Observation 7a589746-4b9d-4003-8223-67cd994012ec · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:55be06a8eca26611bd8cc608aa0fe1ec597e5d94fe2260b5e2370ad44ad73036

Observation 0b3f24fe-3d5a-4d07-8814-eb55a702cb0b · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:3c2265912c8e2ec52c3b894ef53f6b27f8f95dd27a41577ce639d5ead98b0469

Observation 9be0412e-d611-4885-b3b4-fb59c8978883 · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:c8d1943c66d65e6518b525c3cc358156d3c781ec04cf04042feb8f81a6193f46

Observation 4d55759f-67c0-4e2a-b659-afc03e68f1b8 · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:32d209f339cd80dfcd8beea1b2998fc2a3f7426278d7321893d0c53807ea5014

Observation 1e8cf8cf-78d0-46c0-870a-db8fe712ad51 · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:4a5636ea6caf5c5ec9a40cf790190e47f43d7ceb11f3969f1609e338f4eebba3

Observation 711d9049-3b33-4cd5-8bba-73bab9ee1a61 · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:ffdc0154ed3d6adb101f3af4aa46c7030e7cefc35b63545aa943d3ffca5e4989

Observation 9017a485-850b-415f-8045-70c0e3263f34 · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:8590adde1f5758edc6b6d533469462c6d4e4cf490fb6c5a40a745153512f100a

Observation 5975ba5e-8dce-41f6-9808-27977a950522 · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:62d6949e4ccb6e2601545f7ca04c1e0f3cdc5ae343378ba4ceabd0982afdad08

Observation 5172a46a-5c14-4984-b1c0-f2d554865378 · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:ab706002e8fb0d201853af20f1aa7d0759ea7bab50ab0d8ea53ef586c6f93eb4

Observation 790696ad-21be-42dd-bb3d-10d92313f679 · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:00663c55832542eb8992cd0b335254c95476d174f9c95b2c54d785f0948fccd7

Observation f7c79992-5738-494a-bb16-aa3e581e2980 · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:f18976597ccdc72fcb43891827f09a2042bc12f29167a95894aa40af22739ab6

Observation 5d076cd4-016d-457b-890f-2425aff52f11 · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:b30019ac1c8b50409d8302c0eeb6037945ed0c2591c2e7a114cc5d7774012398

Observation 96f61686-b245-40d7-9e0c-581bfb9ea0bb · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:94ffe78eaa45df19fbbcac1862f61d0cbd2ea073fb6e201bbb24edb7cbced650

Observation 8414f0d0-907b-4e98-ba8b-56eab90d50db · outbound

This paper cites an unresolved cited work.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:5f59578d5ba78acab98c70d8c6d0a83ec45c4daed3821d80466b4763635fe6a7

Observation c15dfecf-6b25-46aa-8325-26554c3fe840 · outbound

This paper cites results": [ {.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition results": [ {

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:f15db1661c7dc4f05a9b6b7c47dc0aff1909af460b11460288b78c1ef17f4de2

Observation 92b41ef1-4290-47ed-b83d-6d0fcc6c5856 · outbound

This paper cites subproblem accuracy goes up, therefore root accuracy goes up.

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition subproblem accuracy goes up, therefore root accuracy goes up

Reference 46

Resolution
malformed identifier
no resolver link, observed 2026-06-28T06:52:15.641447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T06:52:15.641447Z digest=sha256:436ac3f5de38abbb5ccc1c4771eff93b2c800e9932cb436cbb9fa55a2e10d2a8

Pith citing papers

No inbound Pith citation observations are available.