Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T05:26:02.449865Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2604.17696.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T05:26:02.449865Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation aab283a5-d4a0-44c9-8390-ed44f5275c40 · outbound
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e4461bb7-6971-4847-9767-edba6d219f73 · outbound
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba915719-2774-496f-9c0f-e50e75e0525e · outbound
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fdc0b025-b4ac-4a67-8597-b42c445674a5 · outbound
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2087c568-e9fa-4960-9428-95892ddd9939 · outbound
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c701ad3-8aba-4837-9f90-3d948c74514f · outbound
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Alternating Turn Structure.In our formulation, players take turns rather than acting simultaneously
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eae4d4f7-66b2-485e-9daa-e869f346855b · outbound
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Each full training run completes in approximately 30 hours
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0930aa1a-80ab-4ff6-a1bc-fd67ec0a49a6 · outbound
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play The game uses only three cards (Jack, Queen, King), where each player receives one card and must de- cide whether to bet, call, or fold based on incom- plete information
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6aec86a1-c33e-4a3a-9bd4-b242fdbf1b63 · outbound
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c13c6e0e-ddb3-493e-9bb0-241bfd5dd0d7 · outbound
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1465eb28-bc40-4be0-b591-425a12016dd8 · outbound
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80c5e8b4-b5db-43f2-8699-b1454f3b08b2 · outbound
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play K.3 Role-Conditioned Advantage Estimation A critical challenge in two-player games is that the expected return differs by role
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42a5510a-22f8-4a0e-958b-896384bd50ca · outbound
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play King beats Queen
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a28b0be5-c6f0-45e1-919e-270b2933c924 · outbound
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.