Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:26:42.128580Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2507.15788.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:26:42.128580Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T17:42:38.122144Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T23:57:28.219193Z
24 of 24 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e8f28327-824b-4bef-a666-a4a54e3e7f69 · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1266b3c4-f9f9-4b9f-adc5-e638084623e2 · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1036b7b4-9e7c-4dc5-911d-4c16ed9ec821 · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Understanding Social Reasoning in Language Models with Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50f9f220-21c6-416f-ba66-10950a797d98 · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aed2f5a8-7297-4374-ac2f-20f013621a3b · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e1dbb0e-a190-4d4b-b6bb-4a1d5eb046ee · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Evaluating Large Language Models in Theory of Mind Tasks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4633569d-2835-40d6-a3c2-a03082a308e6 · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96a35b91-4428-4388-bdb4-5174da70c734 · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 007e2f8e-39e2-475c-8829-2f9866b79efa · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc647f6b-0f7c-41b7-bbc8-412662dd788c · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Training language models to follow instructions with human feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21b2a006-8d70-423c-b673-5d7f54e56a55 · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b57c805-5331-46eb-bbda-e90ebd9232be · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e34b3763-d21e-407d-bba4-35096ec82087 · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f018110-6188-4648-a5d3-ab972d478bf2 · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cbe69e1-c50d-471d-b8ab-7ed36264d93d · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e13adfdd-b3a5-4acc-8178-6319d8ade733 · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a76be41-bfcd-4bbe-906f-37f8621a2200 · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 922a93a1-2cda-4220-8252-7824287ecfff · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a0952fd-084b-4284-bb6d-29914603a735 · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f158a8bf-d79e-4c01-afe9-508d080014c8 · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82e049ef-2111-462a-a260-beab818bf18a · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c78e948d-5acd-45a5-b72d-439dffeffdb8 · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d6c9a4e1-f7f4-4430-b86c-23dbe1170a8c · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning online" 'onlinestring :=
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6a913be-c6a1-4e59-9562-5535a4dbdd76 · outbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning write newline
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 583dd564-c806-4896-918e-5b6e5914a115 · inbound
From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.