Pith. sign in

Paper Citation Record · LEDGER

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning

As of 22 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2507.15788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15788 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:26:42.128580Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T17:42:38.122144Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:57:28.219193Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e8f28327-824b-4bef-a666-a4a54e3e7f69 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 1

Resolution
verified exact
doi, observed 2026-08-06T15:26:42.185748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T15:26:39.695665Z digest=sha256:a75ad213b62ee9241ed9a6df9890d0d46f9d92c5e36daec5b89dafdc3d8bbb08

Observation 1266b3c4-f9f9-4b9f-adc5-e638084623e2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:39.772695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:39.772695Z digest=sha256:1ecb2746c286b4d226813e6f5e142040863c73fee1cbafbc31cc253332321e80

Observation 1036b7b4-9e7c-4dc5-911d-4c16ed9ec821 · outbound

This paper cites Understanding Social Reasoning in Language Models with Language Models.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Understanding Social Reasoning in Language Models with Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:39.871666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:39.871666Z digest=sha256:256e2bb7333a18fada795d033a785179134f07f9f5cf13ef52dd313002be19bb

Observation 50f9f220-21c6-416f-ba66-10950a797d98 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:39.992368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:39.992368Z digest=sha256:3d0c1812124562851af8efb97fb5a4ed7af98f5b69f8ff7e21e0a2a1181000f1

Observation aed2f5a8-7297-4374-ac2f-20f013621a3b · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.154341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.154341Z digest=sha256:2bfbb61f26a97686d65a63ac3e54ea34ea723fcb286fedaefe69e000f531f7c2

Observation 1e1dbb0e-a190-4d4b-b6bb-4a1d5eb046ee · outbound

This paper cites Evaluating Large Language Models in Theory of Mind Tasks.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Evaluating Large Language Models in Theory of Mind Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.278651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.278651Z digest=sha256:b2106796c0dc2c76aeb51d59f71ba737e037a3920661605763be02b465b119b0

Observation 4633569d-2835-40d6-a3c2-a03082a308e6 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.426349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.426349Z digest=sha256:da24aaf9c1ef878410a41f68cc316d1b2935df64cafc9b5c19d07309c7beb2de

Observation 96a35b91-4428-4388-bdb4-5174da70c734 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.526455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.526455Z digest=sha256:75447ae78bc23bf217b9d7691870488d0f93a74c492099d5418918b136ea520d

Observation 007e2f8e-39e2-475c-8829-2f9866b79efa · outbound

This paper cites Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models?.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.632447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.632447Z digest=sha256:cfc2aaf1ef393e41678dbff1ba1a5dc496a285d7a25e34ee0464efc526a9251c

Observation cc647f6b-0f7c-41b7-bbc8-412662dd788c · outbound

This paper cites Training language models to follow instructions with human feedback.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Training language models to follow instructions with human feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.766509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.766509Z digest=sha256:ee248ccb429cfe6cd63fa0127f6919570f650fbf7de21d9446c82f6f32a71f09

Observation 21b2a006-8d70-423c-b673-5d7f54e56a55 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.875592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.875592Z digest=sha256:f054b8f788cfafd2c247d72a56c18907104a1e171cc09dd72fe11c70c33b1985

Observation 0b57c805-5331-46eb-bbda-e90ebd9232be · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:26:42.305311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T15:26:40.974094Z digest=sha256:491808c5249410a8e23fcba87887078a06a0597c3b44520777b48283451c9717

Observation e34b3763-d21e-407d-bba4-35096ec82087 · outbound

This paper cites Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:26:42.220425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T15:26:41.077160Z digest=sha256:6b0984c9f6c3b881f761a89536e844e07a3c4ad25e60fbf68c5ed93ce4aba4c0

Observation 1f018110-6188-4648-a5d3-ab972d478bf2 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.120715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.120715Z digest=sha256:fbe8f6fb1202e520e98a5b4d37fe6f2662f5c5fa99cb5206250171f7b3a15da2

Observation 3cbe69e1-c50d-471d-b8ab-7ed36264d93d · outbound

This paper cites Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.221542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.221542Z digest=sha256:f66bddc6d452158ea080603be206c6b96f50edd554e3bc0b3c2dff4a00c777c0

Observation e13adfdd-b3a5-4acc-8178-6319d8ade733 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.350355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.350355Z digest=sha256:dae35119c26e0149ed8bc1d62b5b34ffc3196897329e0c7f5df6a4de76824a78

Observation 6a76be41-bfcd-4bbe-906f-37f8621a2200 · outbound

This paper cites Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.514564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.514564Z digest=sha256:04e11803da32835f4d46d082dc58a59eec06409914fad7275f7d2e449be2439d

Observation 922a93a1-2cda-4220-8252-7824287ecfff · outbound

This paper cites Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.639155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.639155Z digest=sha256:13131522f5a3bfe902d83444f345bade584df00472b068710637b5dfcaa6d0ab

Observation 6a0952fd-084b-4284-bb6d-29914603a735 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:26:42.298530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T15:26:41.749750Z digest=sha256:205cb855ee7eec4bd321b4db49e654e3dd85374285bc6e88e1ffe871c89f3f12

Observation f158a8bf-d79e-4c01-afe9-508d080014c8 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.829173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.829173Z digest=sha256:20ba0fe4e672e5193f63259c254b4cefe71b6ebd8a1e373141db7ea9eb239e48

Observation 82e049ef-2111-462a-a260-beab818bf18a · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.957990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.957990Z digest=sha256:fcd2c060ddfe66bce252ff7359912ce0ba299054ca4c2480443da0994ddaae32

Observation c78e948d-5acd-45a5-b72d-439dffeffdb8 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:26:42.290762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T15:26:42.039503Z digest=sha256:dce5c0d3f71590e56ac199eba2d62fb36cdec93f9481a5ebd42719c93147392d

Observation d6c9a4e1-f7f4-4430-b86c-23dbe1170a8c · outbound

This paper cites online" 'onlinestring :=.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning online" 'onlinestring :=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:42.116273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:42.116273Z digest=sha256:90a178ecf14bbe53f31e5724f49fced25d23a85a35b818cffd3b11c422329093

Observation e6a913be-c6a1-4e59-9562-5535a4dbdd76 · outbound

This paper cites write newline.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning write newline

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:42.128580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:42.128580Z digest=sha256:b55b99751e3f42d02b9bcfd8fdbce698354569ef13613c250faa60597f97f80d

Pith citing papers

Observation 583dd564-c806-4896-918e-5b6e5914a115 · inbound

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning cites this paper.

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:57:28.220541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T17:42:38.122144Z digest=sha256:cb66c5ec42c285b28ea832aef21443910af250bafb969092af8494c4154a7031