Pith. sign in

Paper Citation Record · LEDGER

Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play

As of 7 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2604.17696.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.17696 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T05:26:02.449865Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved7
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aab283a5-d4a0-44c9-8390-ed44f5275c40 · outbound

This paper cites A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios.

Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:01:49.601685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:26:02.449865Z digest=sha256:160eeaf7ef4831fb72dc86b2d1f4e23cc23ec7787c0fc0dc9f4078147a4a61c1

Observation e4461bb7-6971-4847-9767-edba6d219f73 · outbound

This paper cites Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games.

Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games

Reference 2

Resolution
malformed identifier
local_arxiv, observed 2026-05-10T07:01:49.599143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:26:02.449865Z digest=sha256:0ca56c11b39db30697c7e9a33694253a9d0b1e16feb6ef88b21637617e47dd11

Observation ba915719-2774-496f-9c0f-e50e75e0525e · outbound

This paper cites an unresolved cited work.

Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-21T20:50:37.378708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:26:02.449865Z digest=sha256:8d7dfdd3cf43dea9aa4c10a2b6749be302f5e2be9dba6c062711723d47f5f6f2

Observation fdc0b025-b4ac-4a67-8597-b42c445674a5 · outbound

This paper cites an unresolved cited work.

Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-21T20:50:37.386263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:26:02.449865Z digest=sha256:26d8baaf8021fdf2fc6684a45898d27e14b36d848b6ce11d52e5933a661b01d4

Observation 2087c568-e9fa-4960-9428-95892ddd9939 · outbound

This paper cites an unresolved cited work.

Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-21T20:50:37.376146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:26:02.449865Z digest=sha256:4038702e1a5f6b7efd1ab8fffe7ebbb5a245c7cebcda9e17cc6114b4aee8e746

Observation 8c701ad3-8aba-4837-9f90-3d948c74514f · outbound

This paper cites Alternating Turn Structure.In our formulation, players take turns rather than acting simultaneously.

Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Alternating Turn Structure.In our formulation, players take turns rather than acting simultaneously

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:50:37.373424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:26:02.449865Z digest=sha256:a6106cf6bd3338cae28b2f52520d287473cf3a4a7a089c981079d905399f3fe9

Observation eae4d4f7-66b2-485e-9daa-e869f346855b · outbound

This paper cites Each full training run completes in approximately 30 hours.

Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Each full training run completes in approximately 30 hours

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:50:37.391885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:26:02.449865Z digest=sha256:5d77f19945d3f07ecb4bcb09ed1d0ddb36060d6b6f93ede5c143ec82455578ed

Observation 0930aa1a-80ab-4ff6-a1bc-fd67ec0a49a6 · outbound

This paper cites The game uses only three cards (Jack, Queen, King), where each player receives one card and must de- cide whether to bet, call, or fold based on incom- plete information.

Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play The game uses only three cards (Jack, Queen, King), where each player receives one card and must de- cide whether to bet, call, or fold based on incom- plete information

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:50:37.389057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:26:02.449865Z digest=sha256:96ee665259f37a10d7520d67cb8efe9b79f2d924de2728915fac4597bd34f871

Observation 6aec86a1-c33e-4a3a-9bd4-b242fdbf1b63 · outbound

This paper cites an unresolved cited work.

Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-21T20:50:37.381384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:26:02.449865Z digest=sha256:36c083aadc23f494dcd914c08f67d0b3912c2f2518214a5f2f304ffdd18122be

Observation c13c6e0e-ddb3-493e-9bb0-241bfd5dd0d7 · outbound

This paper cites an unresolved cited work.

Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-21T20:50:37.383782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:26:02.449865Z digest=sha256:ec5f4823eefb05318e88c2677a3c5dbeca5d4e13366b6d8aaa97fb6b736af5a9

Observation 1465eb28-bc40-4be0-b591-425a12016dd8 · outbound

This paper cites an unresolved cited work.

Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-21T20:50:37.370571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:26:02.449865Z digest=sha256:965bb386addeeb0eec9bc8d2ff3e3e217c2d5533421f42929e34c3a23da31704

Observation 80c5e8b4-b5db-43f2-8699-b1454f3b08b2 · outbound

This paper cites K.3 Role-Conditioned Advantage Estimation A critical challenge in two-player games is that the expected return differs by role.

Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play K.3 Role-Conditioned Advantage Estimation A critical challenge in two-player games is that the expected return differs by role

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:50:37.394588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:26:02.449865Z digest=sha256:516ae0f3975937452fb5a618447758565980929ac132d2796b58c84382718cb5

Observation 42a5510a-22f8-4a0e-958b-896384bd50ca · outbound

This paper cites King beats Queen.

Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play King beats Queen

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T20:50:37.397074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:26:02.449865Z digest=sha256:29b7ec632e60391d5250307f463a87104192a87779f193997837f90214d79347

Observation a28b0be5-c6f0-45e1-919e-270b2933c924 · outbound

This paper cites an unresolved cited work.

Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-21T20:50:37.399481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:26:02.449865Z digest=sha256:552cd62511bba98cb98699a4450088391a19626409916b3a2e0f28e7b1221031

Pith citing papers

No inbound Pith citation observations are available.