Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T00:58:20.234462Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2605.08327.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T00:58:20.234462Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T10:39:43.918978Z
A source-named dated measurement, never combined with another source.
Source: cited_works
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a8f166f2-6b35-4f36-bcda-73ca6f082c33 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation Language models are few-shot learners
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8614897c-ff57-436c-948e-d64bb12965b6 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 655173ec-9763-4c41-b3fa-bb08a8178cd1 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation ReAct: Synergizing Reasoning and Acting in Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 51c866c6-4a88-4279-beb3-50f59cd53cd0 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation Training Verifiers to Solve Math Word Problems
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5b08f328-6fd1-4fa2-8133-451c0b909d01 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation Constitutional AI: Harmlessness from AI Feedback
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 50f64533-aa81-4f76-bf10-e038927d9b25 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation Co-evolving agents: Learning from failures as hard negatives
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 63ee066b-45b2-419e-95eb-d22a2a50b08d · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation Language self-play for data-free training
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b380d60c-8d5c-490a-b2f9-deb427a067d3 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8ef3611f-10e0-4627-b608-e980cf3dd294 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation Efficacy of Language Model Self-Play in Non-Zero-Sum Games
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2999a3bd-6eb4-49e3-87dd-ea89fa6d899b · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation org/abs/2603.01213
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 277642bd-ad67-4c03-b1f9-ab1757c88cdd · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation The goal structuring notation–a safety argument notation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bdd41249-88dd-42f3-b381-a115c142b8e5 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 83f3ab9e-2b90-4612-b121-7204d9a8084c · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation TaxCalcBench: Evaluating Frontier Models on the Tax Calculation Task
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ca79f63a-bc55-4f78-a343-96678719fee5 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation Qwen3 Technical Report
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2c56853b-7b1c-4b3d-9b08-a818455217e4 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 15d47839-66c6-4cd3-9609-9d2458071c16 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation Group Sequence Policy Optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation aac73811-644f-48b5-93ff-571c0e86ac05 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation It Takes Two: Your GRPO Is Secretly DPO
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 25cb3376-3eab-433b-ab3f-b252dbe8d50c · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation Debating with More Persuasive LLMs Leads to More Truthful Answers
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2d420f23-3d1c-4f25-84cf-206475aebbbf · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation Proximal Policy Optimization Algorithms
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1a9236eb-31dd-48ba-a27f-c55d2bff4e65 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation Geometry of drifting mdps with path-integral stability certificates
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3c586655-12b3-47e1-b5c1-6e2207c36487 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation Qwen2.5-Coder Technical Report
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3b26b06d-d5f6-49aa-88f5-c97d6a2c7bc5 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation Large Language Models as Agents in Two-Player Games
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c50ff168-9937-4c9a-8b70-30dbe38795e3 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation Game of thought: Robust information seeking with large language models using game theory.arXiv preprint arXiv:2602.01708
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ae56dbb1-8d6e-4328-8651-ef4d041acdf7 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation Structuring value representations via geometric coherence in markov decision processes.arXiv preprint arXiv:2602.02978, 2026c
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bcf32bfc-306a-4317-be1b-e928a24c9b0d · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation Agentic ai for cyber defense: Llm-guided hierarchical multi-agent reinforcement learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3a635cd6-b4aa-40ec-8991-02ea91760983 · outbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1f0e8632-ba8d-4bc4-84e2-963ba3751ff5 · inbound
FlowEdit: Information-Theoretic Control of LLM Reasoning Flows for Ill-posed Problems Involving Conflicts Interactive Critique-Revision Training for Reliable Structured LLM Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.