Pith. sign in

Paper Citation Record · LEDGER

When Can LLMs Learn to Reason with Weak Supervision?

As of 4 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 1 inbound Pith citation observation for arXiv:2604.18574.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.18574 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T04:55:29.315134Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T21:26:33.106406Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ada7fa5-c5eb-44ee-93d6-5ca8138d79f6 · outbound

This paper cites Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps.

When Can LLMs Learn to Reason with Weak Supervision? Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps

Reference 1

Resolution
metadata mismatch
doi, observed 2026-05-10T04:55:33.915543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:55:29.315134Z digest=sha256:7ac342e8bca86d1812cd1b376e083cff6be267befef7dcd589e0aa165acf48cb

Observation 089bf9a3-020d-4017-8d8c-9618a37d3d16 · outbound

This paper cites For each prompt, we sample 16 responses from the policy model.

When Can LLMs Learn to Reason with Weak Supervision? For each prompt, we sample 16 responses from the policy model

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:24:31.434278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:55:29.315134Z digest=sha256:9567a4260b3023d36e52bd3024e365eadd13c8e0f6ea24526d89385c74c4f524

Observation 9138869c-eb12-409a-bbb7-a19727298048 · outbound

This paper cites Mass of sucrose required = 0.2moles× 342g/mole= 68.4g.

When Can LLMs Learn to Reason with Weak Supervision? Mass of sucrose required = 0.2moles× 342g/mole= 68.4g

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:24:31.431609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:55:29.315134Z digest=sha256:df73946b2e5614a7d2f8c2e14b0f1f1c8ea387a86e6559a05cecd94375263ce1

Observation 675817a6-c1ae-4051-8244-09022b715383 · outbound

This paper cites an unresolved cited work.

When Can LLMs Learn to Reason with Weak Supervision? Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-22T01:24:31.424258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:55:29.315134Z digest=sha256:8ebf6c578326ac4f5e3b62dcd2da8fc1ab9aa7b72dde0204a7438ec7adbaeaee

Observation ca2e8748-6b35-44c4-a8c8-fc00e2c24d89 · outbound

This paper cites an unresolved cited work.

When Can LLMs Learn to Reason with Weak Supervision? Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-22T01:24:31.426532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:55:29.315134Z digest=sha256:ecf216ef6aa060a661576d004589e78f971375c0c72be28da59a166514d7ccc3

Observation 8d5ecf1e-741d-44d6-bbe7-6e3a3d82580c · outbound

This paper cites an unresolved cited work.

When Can LLMs Learn to Reason with Weak Supervision? Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-22T01:24:31.428869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:55:29.315134Z digest=sha256:e4c82ef628ff57491a81b47ebeeded9d92a0d4c316c0a78c44dc64940202e613

Observation 60166b37-a77b-4cdf-99b3-ad4c10b9b3e4 · outbound

This paper cites • The second limit evaluates to2 √ 2.

When Can LLMs Learn to Reason with Weak Supervision? • The second limit evaluates to2 √ 2

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:24:31.414392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:55:29.315134Z digest=sha256:19e0ba98fa5550f511df43f44831ed1553d866f1d1bcda5d43f79a795375e863

Observation 95379e60-5d28-4b96-b31c-67cc9ab077bf · outbound

This paper cites an unresolved cited work.

When Can LLMs Learn to Reason with Weak Supervision? Unresolved cited work

Reference 8

Resolution
malformed identifier
raw_fallback, observed 2026-05-22T01:24:31.437107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:55:29.315134Z digest=sha256:7dcf1c7f3d0d1dbbef0cce93f410135e8a0510fd3df65334936d38791d314df8

Observation bc32033e-d8ca-46b2-bd2d-4e140b910e80 · outbound

This paper cites an unresolved cited work.

When Can LLMs Learn to Reason with Weak Supervision? Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-22T01:24:31.417099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:55:29.315134Z digest=sha256:9ea7eee38f339019f5f132df9efc060e3075d6b61441f13aa7a46a3cfa228f49

Observation 4eb2e4d8-a00f-4908-a655-27b473b6d43a · outbound

This paper cites an unresolved cited work.

When Can LLMs Learn to Reason with Weak Supervision? Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-22T01:24:31.422035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:55:29.315134Z digest=sha256:f89aedcbbcc752d1803ac498e986d162529082b77d87c084205b645db03f9803

Observation b9d8ca54-d5dd-431a-89a6-390c8bc1864a · outbound

This paper cites This value is calculated as: 10 5 = 10! 5!5! = 252.

When Can LLMs Learn to Reason with Weak Supervision? This value is calculated as: 10 5 = 10! 5!5! = 252

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:24:31.419576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:55:29.315134Z digest=sha256:3c3a6e61663d77d875eb0d9fa5219985ef171e9371948aaef69544a2ea47bb92

Observation 195649da-7415-4732-956d-3a9de15287d9 · outbound

This paper cites an unresolved cited work.

When Can LLMs Learn to Reason with Weak Supervision? Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-22T01:24:31.409041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:55:29.315134Z digest=sha256:5bca4e81693f1ef5d4ed14fbd28a6ea249378ffbdfa92f838646eb132496ef91

Observation c0788a40-b661-46d0-8837-325dc884719e · outbound

This paper cites Therefore, the probabilityPis: P= Number of favorable outcomes Total number of outcomes = 2 252 = 1 126 So the final answer is 1 126.

When Can LLMs Learn to Reason with Weak Supervision? Therefore, the probabilityPis: P= Number of favorable outcomes Total number of outcomes = 2 252 = 1 126 So the final answer is 1 126

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:24:31.411715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:55:29.315134Z digest=sha256:045e03b4d1716ad5755955b2e5cbad7ecd018f2e3988f4dcee5b20f0a15e0d27

Pith citing papers

Observation 5b0e64e7-b806-40ff-a5b4-d7e81e974fcf · inbound

Understanding Reasoning from Pretraining to Post-Training cites this paper.

Understanding Reasoning from Pretraining to Post-Training When Can LLMs Learn to Reason with Weak Supervision?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T21:26:33.106406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:26:33.106406Z digest=sha256:79f61a8b7d348665ec79439dfb608363698616d0815a30d0fcc380cf2ba159dc