Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-25T21:04:09.237686Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2606.25556.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-25T21:04:09.237686Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a852350c-56eb-4c8f-a697-f7d81fee59d2 · outbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents 3SPO: State-Score-Supervised Policy Optimization for LLM Agents
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 78dd07e1-454d-43d5-a82c-1294cc757c56 · outbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents Hierarchy-of-groups policy optimization for long-horizon agentic tasks
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d11bf8e8-69d3-4f59-9f74-c090727beb5f · outbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents Counterfactual Credit Policy Optimization for Multi-Agent Collaboration
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3836ffba-fb46-4ec7-9a06-ecfcd9361482 · outbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8512bf8a-c8f5-4cb9-b9a4-09104d992594 · outbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents ADaPT: As-needed decomposition and planning with language models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38c871c4-6484-4b15-b2c0-f7ec0c06baef · outbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents Proximal Policy Optimization Algorithms
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ed63e954-a7aa-44b3-bdad-2749b8b4bac1 · outbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2d4e19af-aff9-440c-b889-168c4c22c501 · outbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cc7153b7-09e3-4ffc-a0cb-6bf605b8d134 · outbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents arXiv preprint arXiv:2603.08754 , year=
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 84eb6661-9c47-46ac-8675-9f5b7b1459d8 · outbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents arXiv preprint arXiv:2505.22338 , year=
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 42f73ad9-1957-40da-a546-5dc98e4bd52c · outbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents Reinforcing multi-turn reasoning in llm agents via turn-level credit assignment
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4adb449d-0218-46a0-a110-4838e5e871f2 · outbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents Qwen2.5 Technical Report
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 352d998d-fff4-42ed-a36d-3956d1af82a9 · outbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents Learning Invariant Representations for Reinforcement Learning without Reconstruction
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 89dc4a3d-182d-438f-a3b8-d4fe67dfe03c · outbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f81d9b8-fd43-4ade-9871-aa0d40c06af6 · outbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents For each promptp we sample G trajectories τ (1),
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16ec4b14-d23e-42c8-b276-fe726f82e4b5 · outbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents see Table I.1
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dc4a11c-36da-43a7-be1b-f26b391035e2 · outbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents OnALFWorld,BiPACElowers the single- ton cluster fraction by 9 .3pp and increases mean group size by 1.6×
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb74eda4-80b1-4b18-8e70-9a7bc3ceb053 · outbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents Val/success-rate (binary aggregate, |V|=128) across three seeds: 93 .8%, 92 .2%, 94 .5%; mean ±std = 93.5±1.2% (reported as 93.5 in theAllcolumn of Table 2)
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.