Pith. sign in

Paper Citation Record · LEDGER

Interactive Critique-Revision Training for Reliable Structured LLM Generation

As of 4 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2605.08327.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.08327 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T00:58:20.234462Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T10:39:43.918978Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact16
  • verified fuzzy3
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a8f166f2-6b35-4f36-bcda-73ca6f082c33 · outbound

This paper cites Language models are few-shot learners.

Interactive Critique-Revision Training for Reliable Structured LLM Generation Language models are few-shot learners

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:07:27.885373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:2f87e37a3b36273363603c80011ac988971beee9e83f93bae20a26d4e1a531cf

Observation 8614897c-ff57-436c-948e-d64bb12965b6 · outbound

This paper cites GPT-4 Technical Report.

Interactive Critique-Revision Training for Reliable Structured LLM Generation GPT-4 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:24.559902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:2ff61d88d9e375ac0c7bff35b975c1d58b0c2cf19172d201d78792bedd33f10b

Observation 655173ec-9763-4c41-b3fa-bb08a8178cd1 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Interactive Critique-Revision Training for Reliable Structured LLM Generation ReAct: Synergizing Reasoning and Acting in Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:24.531734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:09ce93d6d3d7bbe0afb7d4be8e47391f9d04ffb6b1e3c24d837dfe0ba3574bd9

Observation 51c866c6-4a88-4279-beb3-50f59cd53cd0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Interactive Critique-Revision Training for Reliable Structured LLM Generation Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:24.635678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:4a1c9c771860aaf386caca8c372f42dec562a4ccda755264f981c40f1d7b7c7b

Observation 5b08f328-6fd1-4fa2-8133-451c0b909d01 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Interactive Critique-Revision Training for Reliable Structured LLM Generation Constitutional AI: Harmlessness from AI Feedback

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:24.537813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:7efe04735e59dbdfdf262b28e3e1d0cead30015e737e7df72fa52a9afb6a98b5

Observation 50f64533-aa81-4f76-bf10-e038927d9b25 · outbound

This paper cites Co-evolving agents: Learning from failures as hard negatives.

Interactive Critique-Revision Training for Reliable Structured LLM Generation Co-evolving agents: Learning from failures as hard negatives

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:36:24.571140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:a4fb4b35eed4c07a288640a56f478e82653e00a795b0134a7c23d59337d563a4

Observation 63ee066b-45b2-419e-95eb-d22a2a50b08d · outbound

This paper cites Language self-play for data-free training.

Interactive Critique-Revision Training for Reliable Structured LLM Generation Language self-play for data-free training

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:36:24.604368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:6212e5788cd8bc954f94c07261953a1cae37f695246fd5a68d02cfeaa66ad254

Observation b380d60c-8d5c-490a-b2f9-deb427a067d3 · outbound

This paper cites Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.

Interactive Critique-Revision Training for Reliable Structured LLM Generation Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:24.548734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:22f5f018c6de18e346bf87a5424523c3f44e270796373b30b664bd6cadf0dc06

Observation 8ef3611f-10e0-4627-b608-e980cf3dd294 · outbound

This paper cites Efficacy of Language Model Self-Play in Non-Zero-Sum Games.

Interactive Critique-Revision Training for Reliable Structured LLM Generation Efficacy of Language Model Self-Play in Non-Zero-Sum Games

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:36:24.615677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:edd6360e13476268ccf369f98da99be6f69e294869f7be33ae63bc6627e1ccad

Observation 2999a3bd-6eb4-49e3-87dd-ea89fa6d899b · outbound

This paper cites org/abs/2603.01213.

Interactive Critique-Revision Training for Reliable Structured LLM Generation org/abs/2603.01213

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:24.576878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:ddaaedcd7eb7532be0c44df685a5a735bf7ec86f45f8847f5e3b9ec0e3e12495

Observation 277642bd-ad67-4c03-b1f9-ab1757c88cdd · outbound

This paper cites The goal structuring notation–a safety argument notation.

Interactive Critique-Revision Training for Reliable Structured LLM Generation The goal structuring notation–a safety argument notation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:07:27.881852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:6355b782dd71bfed3c1ce61544a3cb29b13d7bdffff7d88908474459597da192

Observation bdd41249-88dd-42f3-b381-a115c142b8e5 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Interactive Critique-Revision Training for Reliable Structured LLM Generation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:24.542814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:a3f52413750b44e389891e1a73487a2839024e6679cddfcadc5d3952da2897c9

Observation 83f3ab9e-2b90-4612-b121-7204d9a8084c · outbound

This paper cites TaxCalcBench: Evaluating Frontier Models on the Tax Calculation Task.

Interactive Critique-Revision Training for Reliable Structured LLM Generation TaxCalcBench: Evaluating Frontier Models on the Tax Calculation Task

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:24.651895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:c8868c8f7286cdd06cd3ae6d4a433e117c08ecd18e735e67be948041474454a6

Observation ca79f63a-bc55-4f78-a343-96678719fee5 · outbound

This paper cites Qwen3 Technical Report.

Interactive Critique-Revision Training for Reliable Structured LLM Generation Qwen3 Technical Report

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:24.581598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:3c0a44cb301c1fd4626416b9f7a20a3bc9b34fe93fdc497372475167e2cbfb75

Observation 2c56853b-7b1c-4b3d-9b08-a818455217e4 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Interactive Critique-Revision Training for Reliable Structured LLM Generation DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:24.565036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:bcdf7723f0eaca2a1495685672ca31ee435e732208f544c666e45683d3a35de9

Observation 15d47839-66c6-4cd3-9609-9d2458071c16 · outbound

This paper cites Group Sequence Policy Optimization.

Interactive Critique-Revision Training for Reliable Structured LLM Generation Group Sequence Policy Optimization

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T08:36:24.609483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:71ff49784217ad79d9ab2cd0d1df8bd63d19df94a369b09e59c90b4c78ad63aa

Observation aac73811-644f-48b5-93ff-571c0e86ac05 · outbound

This paper cites It Takes Two: Your GRPO Is Secretly DPO.

Interactive Critique-Revision Training for Reliable Structured LLM Generation It Takes Two: Your GRPO Is Secretly DPO

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:43:11.007781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:2aa5a78948b109a658479481594086f80cc0569bfb616e4201eb4324b0fdebb1

Observation 25cb3376-3eab-433b-ab3f-b252dbe8d50c · outbound

This paper cites Debating with More Persuasive LLMs Leads to More Truthful Answers.

Interactive Critique-Revision Training for Reliable Structured LLM Generation Debating with More Persuasive LLMs Leads to More Truthful Answers

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:24.526604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:4b82508637a00fa7174e874fe5208971ba15b05b75a252dc174ecce21a1725b6

Observation 2d420f23-3d1c-4f25-84cf-206475aebbbf · outbound

This paper cites Proximal Policy Optimization Algorithms.

Interactive Critique-Revision Training for Reliable Structured LLM Generation Proximal Policy Optimization Algorithms

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:24.592786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:0a90e6aee1c99bdeab43607a8bcbc666753a930d30460dcede51f01bb4da9b81

Observation 1a9236eb-31dd-48ba-a27f-c55d2bff4e65 · outbound

This paper cites Geometry of drifting mdps with path-integral stability certificates.

Interactive Critique-Revision Training for Reliable Structured LLM Generation Geometry of drifting mdps with path-integral stability certificates

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:24.626743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:69d137bbd9f58026884028364d405701fd63fa384148777af5dccfb72c0ed608

Observation 3c586655-12b3-47e1-b5c1-6e2207c36487 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Interactive Critique-Revision Training for Reliable Structured LLM Generation Qwen2.5-Coder Technical Report

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:24.553795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:b01448ca59575efe8a0496b1b56450ec0282122777e8a2389fa0eff3a16a3793

Observation 3b26b06d-d5f6-49aa-88f5-c97d6a2c7bc5 · outbound

This paper cites Large Language Models as Agents in Two-Player Games.

Interactive Critique-Revision Training for Reliable Structured LLM Generation Large Language Models as Agents in Two-Player Games

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:36:24.621296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:d788a9b4a1648aaf54fa0df3fce13bf5b92008341cbce15156feab5a09703b7a

Observation c50ff168-9937-4c9a-8b70-30dbe38795e3 · outbound

This paper cites Game of thought: Robust information seeking with large language models using game theory.arXiv preprint arXiv:2602.01708.

Interactive Critique-Revision Training for Reliable Structured LLM Generation Game of thought: Robust information seeking with large language models using game theory.arXiv preprint arXiv:2602.01708

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:24.598462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:a647e8d897067081377c57963352755747b704af0e7bfe757a70a44af180dda5

Observation ae56dbb1-8d6e-4328-8651-ef4d041acdf7 · outbound

This paper cites Structuring value representations via geometric coherence in markov decision processes.arXiv preprint arXiv:2602.02978, 2026c.

Interactive Critique-Revision Training for Reliable Structured LLM Generation Structuring value representations via geometric coherence in markov decision processes.arXiv preprint arXiv:2602.02978, 2026c

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:24.642603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:f9008b0d1e53ac9de7e82a5a7ee66016b4589323f0bc42f08336dbf61814a2f8

Observation bcf32bfc-306a-4317-be1b-e928a24c9b0d · outbound

This paper cites Agentic ai for cyber defense: Llm-guided hierarchical multi-agent reinforcement learning.

Interactive Critique-Revision Training for Reliable Structured LLM Generation Agentic ai for cyber defense: Llm-guided hierarchical multi-agent reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:07:27.883726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:92c690cbadd577f1b2b683b781f23cc1ffe32b873afaac319ebbfa84027264ab

Observation 3a635cd6-b4aa-40ec-8991-02ea91760983 · outbound

This paper cites an unresolved cited work.

Interactive Critique-Revision Training for Reliable Structured LLM Generation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-14T09:07:27.879950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:58:20.234462Z digest=sha256:062b38a610112996e73986f1db5a910c2d9dbbe95e5190df178bdc4e50d2eae1

Pith citing papers

Observation 1f0e8632-ba8d-4bc4-84e2-963ba3751ff5 · inbound

FlowEdit: Information-Theoretic Control of LLM Reasoning Flows for Ill-posed Problems Involving Conflicts cites this paper.

FlowEdit: Information-Theoretic Control of LLM Reasoning Flows for Ill-posed Problems Involving Conflicts Interactive Critique-Revision Training for Reliable Structured LLM Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T10:39:43.918978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:39:43.918978Z digest=sha256:4b8357d8e56ab5250693ead85946d1382ad558e81d3c82f58052166966a74c6e