Pith. sign in

Paper Citation Record · LEDGER

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents

As of 18 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2608.13179.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.13179 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:03:54.167567Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02b2cbbf-c0eb-469d-9164-3e8d22303f8e · outbound

This paper cites Are: Scaling up agent environments and evaluations.arXiv preprint arXiv:2509.17158,.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Are: Scaling up agent environments and evaluations.arXiv preprint arXiv:2509.17158,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.097497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.097497Z digest=sha256:4486980fef32c5138a944e8736aaf41e85f8dc01256a03aa7b9ad2c683545190

Observation 6a615f59-ff61-47ab-990b-9e9c0f0f2b58 · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.100378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.100378Z digest=sha256:f1184279c8fbfa845c79bc2741b3056abab227de8d1ac3c54fd7b89cca52c967

Observation 20da41db-3010-49b0-b63d-7e7a095a20f4 · outbound

This paper cites GPT-4o System Card.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents GPT-4o System Card

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.103487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.103487Z digest=sha256:5a029a9109bb98696b7294ff19d16e6dc43f1137460e8bacc508d7533989bafe

Observation c5a996b9-8daa-4a85-9d07-d38cd9762048 · outbound

This paper cites Swe-bench: Can language models resolve real-world github issues? InInternational Conference on Learning Representations, volume 2024, pp.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Swe-bench: Can language models resolve real-world github issues? InInternational Conference on Learning Representations, volume 2024, pp

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.106490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.106490Z digest=sha256:b51792493e03b370201a2939e15d56971a81aa267f447c8eab8fcb5d651a51ad

Observation 519b9f45-a30b-4fee-add1-f051420339df · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.109252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.109252Z digest=sha256:516511aa42513de31e773e8ae7c5478baa63a8434d496ed878f15ac00481d6d4

Observation 7539385c-ba54-4a52-b65a-eede7b3ce3db · outbound

This paper cites Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.112293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.112293Z digest=sha256:aa4e3f9b3f31fb2070f6e2b808afe23e1a716b21209a32659b829d02c0f5a7d2

Observation e115788a-ad17-4f6b-93d7-e7d2bf47debb · outbound

This paper cites LLMs Get Lost In Multi-Turn Conversation.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents LLMs Get Lost In Multi-Turn Conversation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.115247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.115247Z digest=sha256:9a26e6fc0af3a90e3ebfb4f0d27319d672752582bbdcd0bde2a0b4d77eeb4ce2

Observation 4eb695ef-12eb-4f0c-a6c7-772a2adaa3c9 · outbound

This paper cites Api-bank: A comprehensive benchmark for tool-augmented llms.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Api-bank: A comprehensive benchmark for tool-augmented llms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.118102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.118102Z digest=sha256:4b689b9ba2114d22a91323a7524813f0d48f4df8cb9246420bd5879bf8d21df8

Observation 1237f7ba-26f3-42e5-9d88-dd6ffda5afd8 · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.122957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.122957Z digest=sha256:33674f85ffbe1717d88a9982dd6db8a06008e925a2966eafe66b47152db47619

Observation d24829ce-eca8-4da6-922b-5790f2625743 · outbound

This paper cites DeepSeek-V3 Technical Report.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents DeepSeek-V3 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.125445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.125445Z digest=sha256:921152df4af7ba387dcd9e21254ad796123d871e41a86a3337533e50cd32a146

Observation 950a9fe2-2ddb-4643-8921-305b5e4d1575 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Understanding R1-Zero-Like Training: A Critical Perspective

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.127601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.127601Z digest=sha256:b159ee3e17a9487288b76fba669e02e734acecb2faa527f46522b889807cf6bd

Observation 737005be-e9db-437f-8ad8-8aefd54f4536 · outbound

This paper cites Self-Distilled Agentic Reinforcement Learning.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Self-Distilled Agentic Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.130612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.130612Z digest=sha256:634f162280a867bdbc1c775ce2da7fa66cc61d73e46df738bfebc9208836eec6

Observation 2f342420-deba-46fa-91f4-b6f4bce98fe7 · outbound

This paper cites Shishir G Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E Gonzalez.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Shishir G Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E Gonzalez

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:03:54.473551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:03:54.133537Z digest=sha256:281863b9287220c391bbd8a2d98cd545440648518cb02938de7a08674fd8daaf

Observation c0d6e633-0499-4ac2-b14a-8fe661d1ea9b · outbound

This paper cites Privileged Information Distillation for Language Models.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Privileged Information Distillation for Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.136223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.136223Z digest=sha256:19d272a3750ed4923a7c81e39b28a63e3daf8e69657365b6a852feb6bc60e118

Observation a28fe76a-9274-449e-94b6-58c5b8162465 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Proximal Policy Optimization Algorithms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.139516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.139516Z digest=sha256:ef2e4368155d4df22590a0ff1b5c8a781d30b2633c8239d44c47b446e845373f

Observation a34fd52e-7264-4a1f-8170-953ddf0de8c0 · outbound

This paper cites Self-Distillation Enables Continual Learning.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Self-Distillation Enables Continual Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.144898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.144898Z digest=sha256:cbb95c8b6df07f860532a9a62ea51600e848e8ef623da337bac22d23111134db

Observation e1475b65-f187-4343-a66b-46866538396a · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.150674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.150674Z digest=sha256:e082e23f5f4cd8c9c3d057ba2872ebbf8a016b2b84033a0b570fb89a4a9e9d34

Observation fbf3de51-abcd-46ba-bfc3-df2facadb335 · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.153463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.153463Z digest=sha256:0a86e2e7cecb999c5064a21ace8c658c91c10cafedcd4c80f9d279d1e9d5e0c3

Observation a226b292-6974-4760-a81a-3e41fa112902 · outbound

This paper cites Qwen3 Technical Report.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Qwen3 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.156353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.156353Z digest=sha256:19cd38c2f95af9bc05f174a98c105b30592383987ac25ef8815b7823d26b29a7

Observation 674d69ae-b3e4-4b39-9d4a-cc41d56582a6 · outbound

This paper cites Self-Distilled RLVR.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Self-Distilled RLVR

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.159010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.159010Z digest=sha256:1f9d329e99d8d0cb76bdf580ecd062366afbba5cff67528a7e7ca31361b17127

Observation 7ff4c1bc-fd62-4dd1-a8fc-b299bc0ae9d3 · outbound

This paper cites Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.161918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.161918Z digest=sha256:ea7ad96200bfe41c85a034ff5c92ffef60abbe86fedb5f47630a75151d52c126

Observation 1a3f1f54-452a-4f88-a1b6-5baf7729a362 · outbound

This paper cites Deep- researcher: Scaling deep research via reinforcement learning in real-world environments.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Deep- researcher: Scaling deep research via reinforcement learning in real-world environments

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:03:54.465686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:03:54.164742Z digest=sha256:a3cbc1c999f3dd649155ea0afae8ba2c04665f5452ad89d5d4c7e45f1a430083

Observation f9481f4a-55f4-4ea8-a033-134b1fb8b362 · outbound

This paper cites an unresolved cited work.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:03:54.457533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:03:54.167567Z digest=sha256:70517e3e5b2ccf116bb85ed9dbe77e1a13c1368fb40becb93d76d66841003339

Observation 22fcd09d-8590-478e-bd34-f051e9696464 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.142171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.142171Z digest=sha256:0c16b2ad718075318315c8d4687fca9729bb340027dcf084ebe0cc3dc7c2110d

Observation ede1e056-57c1-40ff-b0bf-a328c1ba7887 · outbound

This paper cites A Survey of On-Policy Distillation for Large Language Models.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents A Survey of On-Policy Distillation for Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.147655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.147655Z digest=sha256:84a610d81e800dc0ddc9733dbf5ef0a229ba2b18b8b49a720d90b78400960ffc

Observation cfeb87f9-c46e-4edb-8262-dfd60b8b77e8 · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents ToRL: Scaling Tool-Integrated RL

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.120450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.120450Z digest=sha256:ea83dd8da7d3283f730ac75ada0089008eb5a087532f76ed3b74ade6ebb48f7e

Observation db0dd0d9-f8c6-4cf1-943c-2577e7c50b27 · outbound

This paper cites On SFT, RL, and on-policy distillation.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents On SFT, RL, and on-policy distillation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.087229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.087229Z digest=sha256:19f14e512a88636b8ebc7d075bd615ef3ce8bbf959404e1bc0f138f5fdd3ea0d

Observation e687b818-9da8-4dcf-ba2c-fded15d60dc2 · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.094279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.094279Z digest=sha256:c34132acc74abf21c490c0aa56c1e3051e133b73ea6871ce3d284de797309bc8

Observation 99861ed0-effa-4125-b844-2b26c22cdd62 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.090793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.090793Z digest=sha256:6df870969bb1843480474603f8ba14ae5fa85bd8705ad18c68860c3cd8ef5289

Pith citing papers

No inbound Pith citation observations are available.