Pith. sign in

Paper Citation Record · LEDGER

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2507.15788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15788 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:26:42.128580Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T17:42:38.122144Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:57:28.219193Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e8f28327-824b-4bef-a666-a4a54e3e7f69 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 1

Resolution
verified exact
doi, observed 2026-08-06T15:26:42.185748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:26:39.695665Z digest=sha256:7c1e59719b45e1b15d736e21dd6c093a3c48f8f5a98619247b09109378b70bd6

Observation 1266b3c4-f9f9-4b9f-adc5-e638084623e2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:39.772695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:39.772695Z digest=sha256:f79b51ffb8e64b4df13b9918224e0ba62da9209bb5b399185ec80e8b1da1aa76

Observation 1036b7b4-9e7c-4dc5-911d-4c16ed9ec821 · outbound

This paper cites Understanding Social Reasoning in Language Models with Language Models.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Understanding Social Reasoning in Language Models with Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:39.871666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:39.871666Z digest=sha256:86c9bce06133e0597ef8fa180ec4c16f8ae11b4e864df13a1d7052ffd4116430

Observation 50f9f220-21c6-416f-ba66-10950a797d98 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:39.992368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:39.992368Z digest=sha256:067050f504457608f53502c7eb0d6d5dd40ea88739c6899b9c3956b96129f3a3

Observation aed2f5a8-7297-4374-ac2f-20f013621a3b · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.154341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.154341Z digest=sha256:adfe02f9cafd059141f535aa97d1a4648044c73704072065516d03c645ff28f0

Observation 1e1dbb0e-a190-4d4b-b6bb-4a1d5eb046ee · outbound

This paper cites Evaluating Large Language Models in Theory of Mind Tasks.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Evaluating Large Language Models in Theory of Mind Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.278651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.278651Z digest=sha256:93d5e025569fe3b3e77e8036d3a6c8dcbfa56a150fe78b0d485b960c1ab1405c

Observation 4633569d-2835-40d6-a3c2-a03082a308e6 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.426349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.426349Z digest=sha256:b909079fa96c0a434c1b4991fae2b12e05630c5d30f6194c78b5f8b23fc0d709

Observation 96a35b91-4428-4388-bdb4-5174da70c734 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.526455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.526455Z digest=sha256:fa199fb6c2370904865fe0a93ae767f2a2c6f24cb1c64ec402ed473ffedc5292

Observation 007e2f8e-39e2-475c-8829-2f9866b79efa · outbound

This paper cites Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models?.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.632447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.632447Z digest=sha256:1c80e7707f06eab90161ab7d9feda952e810e94ccb892b25a497098604a0dc9c

Observation cc647f6b-0f7c-41b7-bbc8-412662dd788c · outbound

This paper cites Training language models to follow instructions with human feedback.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Training language models to follow instructions with human feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.766509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.766509Z digest=sha256:b4ad1c6683881e814f7f0de02200521e49d5effc5a918f6b0010c6f1ea703e10

Observation 21b2a006-8d70-423c-b673-5d7f54e56a55 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.875592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.875592Z digest=sha256:e6092206a78efeac1b65496d5df82205d2c180c08e70a964b6f5167566070f7f

Observation 0b57c805-5331-46eb-bbda-e90ebd9232be · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:26:42.305311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:26:40.974094Z digest=sha256:351cc59468fbb9287c89dc51254f552ffcbbdb4e712115257eb382ddde4910c9

Observation e34b3763-d21e-407d-bba4-35096ec82087 · outbound

This paper cites Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:26:42.220425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:26:41.077160Z digest=sha256:6629bd2573d45b88ef8b939ef518549220fef10adf475eb6aa4e1eec765f29e9

Observation 1f018110-6188-4648-a5d3-ab972d478bf2 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.120715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.120715Z digest=sha256:dc1a0c51274a99de1cf7be28a6649d700455e15aa5869de69f6c2ea64fccd7e9

Observation 3cbe69e1-c50d-471d-b8ab-7ed36264d93d · outbound

This paper cites Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.221542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.221542Z digest=sha256:11fa5ab582e32a32558b30df7949cd0489a39fbece6671f7ae4bcc449b64bc8f

Observation e13adfdd-b3a5-4acc-8178-6319d8ade733 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.350355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.350355Z digest=sha256:e63dab37b5b9891f329afde2904cbeeef1b2e2233d49be1a3e493ea146a3ad41

Observation 6a76be41-bfcd-4bbe-906f-37f8621a2200 · outbound

This paper cites Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.514564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.514564Z digest=sha256:91ecd64bb31510226ff05f5be9e310e67cbfbbf2125e043c09f46c018ea40309

Observation 922a93a1-2cda-4220-8252-7824287ecfff · outbound

This paper cites Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.639155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.639155Z digest=sha256:7527302ccb00db3f881f923e0ac8f3cb91cc5ba12b8c93435130da9e679911ce

Observation 6a0952fd-084b-4284-bb6d-29914603a735 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:26:42.298530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:26:41.749750Z digest=sha256:302001e2cb9778d7827acf0e6b46acd8fcb15a588d23671565e2942653f9a07d

Observation f158a8bf-d79e-4c01-afe9-508d080014c8 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.829173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.829173Z digest=sha256:f5f72f16cd585d6b10a36497face2b05b6342c068a5838c5bfe8df0e2aaab1c9

Observation 82e049ef-2111-462a-a260-beab818bf18a · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.957990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.957990Z digest=sha256:e6fa807af85ae432c5733e39e76b7972e2085a6feb1a38ec62306a26fd652cd0

Observation c78e948d-5acd-45a5-b72d-439dffeffdb8 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:26:42.290762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T15:26:42.039503Z digest=sha256:f41293ee3d122d422ca17d137c463c287a03a4598f492314e3a014c538f8d6eb

Observation d6c9a4e1-f7f4-4430-b86c-23dbe1170a8c · outbound

This paper cites online" 'onlinestring :=.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning online" 'onlinestring :=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:42.116273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:42.116273Z digest=sha256:016494417434141700bee110a48b5050e81208e8f51eff0652dd33dd46b05b44

Observation e6a913be-c6a1-4e59-9562-5535a4dbdd76 · outbound

This paper cites write newline.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning write newline

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:42.128580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:42.128580Z digest=sha256:cc8997feff9988e00a4412ddb5566a190af22d71363985c6eefe35322e42c024

Pith citing papers

Observation 583dd564-c806-4896-918e-5b6e5914a115 · inbound

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning cites this paper.

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:57:28.220541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T17:42:38.122144Z digest=sha256:1f662e3be76e299432da2a3e85dfc3dd3362978a38d2cc7f2f29c1ed494c2a00