Pith. sign in

Paper Citation Record · LEDGER

State2State: Environment-Derived Mid-Training for LLM Agents

As of 18 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2608.04934.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04934 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:40:40.747744Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bc41c385-6f6e-45ac-9c27-505fb56a10a5 · outbound

This paper cites Tongyi DeepResearch Technical Report.

State2State: Environment-Derived Mid-Training for LLM Agents Tongyi DeepResearch Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:40:40.696250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:40:40.696250Z digest=sha256:1019347d1288bb663a503f146bee16cabffb2ea9abeed76892fccab28cdb9d2f

Observation f8d34a17-fff4-44e3-9407-bdd10225ef3b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

State2State: Environment-Derived Mid-Training for LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:40:40.700383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:40:40.700383Z digest=sha256:3d569fe59c798823817949c99b774dbf99f4484265da372bff787ec750cc011d

Observation 3861aa21-7c54-4201-ad20-470a8422aa3d · outbound

This paper cites InProceedings of the 2025 Con- ference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, pages 3062–3077.

State2State: Environment-Derived Mid-Training for LLM Agents InProceedings of the 2025 Con- ference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, pages 3062–3077

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:40:40.708919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:40:40.708919Z digest=sha256:e67da2a6e8f6d423737feafbb1f95b2de216047c1fa606d51cb80269ebcef6ef

Observation 0724f0aa-8893-44af-aac7-f7f83ed2056c · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

State2State: Environment-Derived Mid-Training for LLM Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:40:40.713363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:40:40.713363Z digest=sha256:694acd980343a4e6625cb242b01e0612526d03bf4a42d2ee5652f75e277ba72b

Observation 9a8f32d7-f285-45a4-8f2c-e24f7d06f64b · outbound

This paper cites action":.

State2State: Environment-Derived Mid-Training for LLM Agents action":

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:40:40.950891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:40:40.734300Z digest=sha256:333f35acb12e21159a2d796c601242a362656b37b9d72752dec5374be4d85e3b

Observation 40f084a9-40bb-4e2c-be14-a26ce8e07740 · outbound

This paper cites an unresolved cited work.

State2State: Environment-Derived Mid-Training for LLM Agents Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:40:40.975416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:40:40.738750Z digest=sha256:a878046ddc7014c92941a9d91671f13d2d2cc611f1872c3a7d1bfc587c638f3b

Observation d0910b67-e55f-4904-b2ee-ee43266a8ae8 · outbound

This paper cites an unresolved cited work.

State2State: Environment-Derived Mid-Training for LLM Agents Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:40:40.963226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:40:40.742961Z digest=sha256:c30099c448faadef016cdc77b008a71140bc4b756d39484ac06f5bc309d0f57b

Observation 63362e1e-7064-48c6-88f6-99ab04cd2bb6 · outbound

This paper cites an unresolved cited work.

State2State: Environment-Derived Mid-Training for LLM Agents Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:40:40.937890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:40:40.747744Z digest=sha256:aa18a2b8577067879846470b0b4982d08152ae2e64d1fce7c0d4831c12356a36

Observation 8892b842-9d66-4b21-81b7-05dd321e4cd1 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

State2State: Environment-Derived Mid-Training for LLM Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 507

Resolution
unresolved
no resolver link, observed 2026-08-15T14:40:40.691781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:40:40.691781Z digest=sha256:62da8646eeb2a96df5fedbbc17006d62075d89ee18d0e0fb377b9017b5a97369

Observation 4a65e4dc-bf5c-406a-abc7-93b0e57f9905 · outbound

This paper cites OpenAI GPT-5 System Card.

State2State: Environment-Derived Mid-Training for LLM Agents OpenAI GPT-5 System Card

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T14:40:40.704545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:40:40.704545Z digest=sha256:ddb5dc0f05dedf768fa9189fa5168bc063fce85cdf0b69537850215e1b4f4904

Observation 6a8c715e-901a-4306-ac77-8a187d70b1ab · outbound

This paper cites action":.

State2State: Environment-Derived Mid-Training for LLM Agents action":

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:40:40.988210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:40:40.722061Z digest=sha256:39378c8d1ff366201bc5a4cc8ffc8180907fe454c130643fddfdd252470d384f

Observation e857b4fd-b8e9-4c3d-9b9b-dde89460a76d · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

State2State: Environment-Derived Mid-Training for LLM Agents Group-in-Group Policy Optimization for LLM Agent Training

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T14:40:40.686820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:40:40.686820Z digest=sha256:c48d9ee5e44f2d85472794fd059c1c5957bd27a735bcbed6ed2a0c1a1500daf7

Observation 09c0c3c2-cc7e-4e5c-9b72-92da381eef3d · outbound

This paper cites an unresolved cited work.

State2State: Environment-Derived Mid-Training for LLM Agents Unresolved cited work

Reference 3077

Resolution
unresolved
no resolver link, observed 2026-08-15T14:40:40.717476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:40:40.717476Z digest=sha256:4bd174ef9f168c0272e9c149b00191dcd8fbf44b2f91b7ed2920f4e488f48f27

Pith citing papers

No inbound Pith citation observations are available.